What are AI evals, and what should leaders know about them?
An AI eval is a repeatable test of whether an AI system does its job well enough: a set of representative inputs, a definition of a good answer, and a way to score the output, rerun as the prompt, model or data changes. For leaders, evals are how you find out a feature fails before customers do, and the people who know the business have to help define what good means.
Level 3: Leading with AI · 3.4 Evals for leaders
AI Evaluation & ReliabilityEnterprise AI
Why AI needs a different kind of test
Ordinary software tests usually check that the same input gives the same, correct output. AI breaks both halves of that. The same input can produce different outputs, and many outputs aren’t simply right or wrong: a summary can be accurate but miss the point, or a reply can be correct but wrong for this customer. You need a way to judge quality across many cases, again and again, as the system changes.
And the gap is widening. Alex Ratner, co-founder and CEO of Snorkel AI, calls it the evaluation gap: “the fact that AI capabilities have been advancing so rapidly beyond our capability to measure them.” He adds that model abilities are jagged. Models now do well on competition math and coding problems, then stumble on mundane tasks with messy inputs and many steps.
What an eval is made of
An eval has four parts:
- Representative inputs: questions, documents or tasks like the ones the system will really see, including real examples and cases it has failed on before. Synthetic cases can fill gaps where real ones are scarce.
- A definition of good: what a correct, useful answer looks like for each input, written as a reference answer or a scoring rubric.
- A scorer: code for anything checkable (did it return valid JSON, cite a document that exists in the source set, stay under a length limit), a second model acting as judge for judgment calls, and people for the cases that matter most.
- A bar and a habit: the results the system has to hit to ship, and the discipline of rerunning the automated checks on every change to a prompt, model or data source.
Hamel Husain, who teaches evals to engineers and product managers, describes three levels in his 2024 essay on evals: quick automated tests run on every code change, human and model review on a regular cadence, and A/B tests with real users after significant product changes. Each level costs more than the one before, which is why they run at different frequencies.
Start by reading real outputs
The most common mistake is starting with a dashboard of generic metrics. Husain’s advice in episode 37 is to “ground it in your failures. So how do you know what your failures are?” His answer is error analysis: read real outputs, label what went wrong, count the failure types, and fix the biggest ones first. He estimates that most of the work of evals is this kind of looking at data. In his field guide he describes teams celebrating a 10% rise in a “helpfulness score” while their users still struggled with basic tasks.
Business people have to define “good”
This is where leaders come in. Loïc Houssier, CTO of Superhuman, told Conor in episode 59, “I think it’s the role of the PM to identify the dimensions” an eval has to cover. His example is a simple question to an email assistant: how much time did you spend in Waymo last month? Answering it means finding the right emails among receipts and marketing, working out what “last month” means relative to today, reading a duration from each trip, and adding them up. Each of those is a dimension to test. Superhuman’s quality bar also comes from the business: anything that makes a user look foolish, like archiving the email that says there’s no daycare tomorrow, is filed as a top-priority bug.
Vikram Chatterji, co-founder of the evaluation company Galileo, sees the same pattern with its customers. In episode 46 he said the LLM judges they use “are actually created not by the developer. They’re created by subject matter experts,” the product managers and specialists who understand the business use case.
Benchmarks aren’t your evals
Public benchmarks and leaderboards tell you how models compare on someone else’s test. Ratner, whose company builds benchmarks, calls them critical, but says “they just obviously can’t be the only tool.” You also need use-case-specific, private tests. Observability, evaluation and benchmarking answer different questions at different stages.
Questions to ask your team
- What are the most common ways this system fails, and how often does each happen?
- Where did the test cases come from, and do they include real failures?
- Who decided what a good answer is, and did someone who knows the business sign off?
- Does the eval run on every change, and what score blocks a release?
- When something fails in production, how does it get into the test set?
If the answers are vague, the team is shipping on demos. The hands-on guide to testing an AI system is the next step down for the people doing the work.
Hear it from the guest
“Ground it in your failures. So how do you know what your failures are?”
“I think it's the role of the PM to identify the dimensions”
“The fact that AI capabilities have been advancing so rapidly beyond our capability to measure them”
Quotes lightly edited to remove filler words.
Go deeper
- Your AI Product Needs EvalsThe widely cited case for evals as the core of AI product work, with three levels of testing and a worked example.
- A Field Guide to Rapidly Improving AI ProductsHow teams that improve fastest work: error analysis, simple data viewers, domain experts writing prompts, and roadmaps that count experiments rather than features.
- AI EngineeringChip Huyen's book on building applications with foundation models. Evaluation gets two chapters early on, ahead of prompt engineering, RAG and agents.
Common questions
- What's the difference between an eval and a benchmark?
- A benchmark is a standardized evaluation used to compare models, like an exam everyone sits. The public benchmarks behind leaderboards test shared, published tasks; the evals that decide whether to ship are tailored to your task, your data and your definition of a good answer. A model that tops a leaderboard can still fail your eval, so public benchmarks help you shortlist models and your own evals tell you whether to ship.
- Can an AI model grade the outputs?
- Yes, and it is how teams score open-ended answers at volume. The method is called LLM-as-a-judge: a second model scores each output against a rubric. It is only trustworthy after you check its scores against human judgments on your own data, and high-stakes calls should still go to a person.
- When should a team start writing evals?
- Before launch, and from real failures rather than a guessed list of metrics. Hamel Husain's advice is to start by reading real outputs, sort the failures into types, count them, and write evals for the types that matter most. The eval set then grows every time production turns up a new kind of failure.
From the conversation
This explainer is drawn from these episodes — each carries its full transcript.