A classifier is not a classifier is not a classifier
Four levels of yes/no decision problems—and how to choose between rules, classifiers, language models, and agents.
From the team
Notes on building reliable AI systems, running models at scale, and the experiments that shape our work.
Four levels of yes/no decision problems—and how to choose between rules, classifiers, language models, and agents.
How reflective optimization aligned eleven models to Sutro’s lead-scoring judgment using only 30 annotations.
How language models are changing synthetic data—and what that unlocks for training, retrieval, simulation, and privacy.
What Google’s Gemini Flash price increase says about the economics of inference and the limits of ever-cheaper intelligence.
Why bulk AI workloads are often cleaner, less expensive, and faster to operate with batch APIs instead of synchronous inference.
A practical comparison of model performance, inference cost, and open-source alternatives for common AI workloads.
A method for seeding diverse language-model outputs, with an accompanying open dataset of one million synthetic people.
What we learned by using small language models to classify more than 40 million Hacker News posts and 10.7 billion tokens.
A practical framework for probing open-source coding models for malicious behavior before bringing them inside your systems.