How to Write Evals for an AI Feature Before You Build It
Last updated September 2026
Writing LLM evals before building means converting the feature’s success criteria into 20 to 40 graded test cases before a single prompt exists. You draw the cases from artifacts you already have: support tickets, search logs, and work your team does by hand today. The eval set becomes the specification, sets the ship bar, and catches every regression after launch.
What does it actually mean to write LLM evals before building?
Almost every eval guide published in the last two years starts in the same place: collect production traces, read them, and group the failures. That advice is correct and it is useless to you right now, because before you build there are no traces. Zero users, zero logs, nothing to annotate.
So the pre-build eval is a different artifact. It is a written specification with a grader attached. You state what the feature must get right, you write the inputs that test it, and you record the output you would accept. Anthropic’s own documentation puts defining test cases as step one of the prompt engineering cycle, before the preliminary prompt and well before any refinement. The order is not stylistic. A prompt written first becomes the definition of correct behavior by default, and every later argument about quality turns into opinion.
This is also the cheapest moment to discover that the feature is a bad idea. Half the AI features that die in review die because nobody could say what a good answer looked like. Writing 30 test cases takes an afternoon and answers that question before anyone spends a sprint on retrieval infrastructure.
Where do test cases come from when nothing is in production yet?
From the work the feature is meant to replace. Every AI feature automates a judgment a human is making today, usually in a system that already keeps records. Six sources, in the order they pay off:
- Support tickets and sales email. Real user phrasing, including the confused and abusive kind. Pull the last 50 and keep the 20 that touch the feature.
- Search and query logs. What people typed when they wanted this and did not get it. The empty result set is your test case list.
- The manual work product. If an analyst writes these summaries today, their last 20 summaries are your reference answers, already graded by the fact that they shipped.
- Your own domain expert, for one hour. Ask them to write ten inputs they expect the system to fail. They will be right about most of them.
- Known adversarial shapes. Empty input, another language, a prompt injection attempt, a question outside scope, an input three times longer than expected.
- Competitor output. Run five real inputs through the shipped competing product and record where it is weak. That is your differentiation, stated as a test.
Group what you collect into five to ten failure themes rather than a flat list. Themes are what you will report on later, and a theme with only one test case in it is a theme you have not thought about yet.
How many test cases do LLM evals before building actually need?
Fewer than the tooling vendors imply, and more than the one example in your ticket. Anthropic’s guidance is blunt on the tradeoff: prioritize volume over quality, because more questions with slightly noisier automated grading beats a handful of carefully hand-graded ones. Its worked examples run to 1,000 cases for sentiment classification, 500 for privacy preservation, 200 for summarization, and 50 for FAQ consistency, which tells you the number is set by how varied the input is, not by how important the feature is.
| Stage | Cases | Who writes them | Grader | What it buys you |
|---|---|---|---|---|
| Before any code | 20 to 40 | Engineer plus one domain expert | Human, by eye | Agreement on what correct means |
| First working prototype | 50 to 100 | Engineer, from the six sources above | Code based assertions | A pass rate you can compare across prompts |
| Before launch | 100 to 300 | Engineer plus expert review | Code based, plus a judge on subjective checks | A defensible ship or no ship decision |
| Weekly in production | 10 to 20 fresh traces | Whoever owns the feature | Human review of outliers | Drift caught in days, not quarters |
Those last two rows come straight from practitioner guidance rather than theory. In his AI Evals FAQ, Hamel Husain recommends reviewing at least 100 traces during error analysis, annotating the first 30 yourself, and then reviewing 10 to 20 traces weekly focusing on outliers. He also gives the most useful sanity check in the field: if you are passing 100 percent of your evals, the eval is too easy, and a 70 percent pass rate usually means the set is actually measuring something.
How do you set a pass bar instead of judging by eye?
By writing the number down before you see the output. A criterion is only measurable if it names a metric, a threshold, and a sample size. OpenAI’s evaluation best practices guide demonstrates the shape with two concrete examples: a transcript summarizer held to a ROUGE-L score of at least 0.40 and a coherence score of at least 80 percent against 1,000 reference transcripts, and a question answering system held to context recall of at least 0.85, context precision above 0.7, and 70 percent or better positively rated answers.
Copy the form, not the numbers. Your version looks like this:
| Dimension | How it is graded | Example bar to set before building |
|---|---|---|
| Task fidelity | Exact match or string assertion | 95 percent of 100 extraction cases return the correct date format |
| Grounding | Assertion that every claim cites a retrieved chunk | Zero uncited claims across 50 cases |
| Refusal behavior | Code check for the refusal string | 100 percent refusal on 20 out of scope inputs |
| Safety | Content filter flag rate | Under 0.1 percent of 10,000 trials flagged |
| Latency | P95 measured in the harness | Under 3 seconds at the ninety fifth percentile |
| Cost | Tokens per answer times price | Under 4 cents per answer at expected context length |
Two of those rows are usually missing from eval sets written after the fact, and they are the two that decide whether the feature survives. Latency and cost per answer are product constraints, not engineering details, and a feature that is accurate at 9 seconds and 30 cents a call has failed. The same discipline applies to the retrieval layer, where most of the surprises live. Our write up of RAG production failure modes covers the ones worth writing a test for on day one, including stale indexes and silent truncation.
When should you add an LLM judge?
After you have labels, not before. An LLM judge is a model you have to evaluate, which means it is only as good as the human labels you validated it against. Husain’s guidance is to label 100 to 200 examples per failure mode, split them into development and test sets with 30 to 50 passes and 30 to 50 fails in each, and then measure the judge’s true positive and true negative rate against those human labels. None of that is available to you the week before you build.
So sequence it. Code based assertions first, because they are deterministic and free. Human spot checks second, on a schedule rather than a whim. The judge third, once real traces exist and you have labeled enough of them to know whether it agrees with you. And when you do build one, ask it for a binary pass or fail with a written critique rather than a score out of five. Nobody, human or model, applies a 1 to 5 scale consistently across a thousand outputs.
What does writing evals before building cost?
More than teams expect, which is exactly why most skip it. Husain reports spending 60 to 80 percent of development time on error analysis and evaluation rather than on building automated checks, and offers a minimum viable version for teams who cannot commit that: 30 minutes manually reviewing 20 to 50 outputs whenever you make a significant change.
Weigh that against the alternative, which is measurable. LangChain’s State of Agent Engineering survey, fielded to 1,340 respondents between 18 November and 2 December 2025 and published in June 2026, found 52.4 percent of organizations running offline evaluations on test sets, 37.3 percent running online evaluations in production, and 29.5 percent not evaluating at all. Among teams with agents actually in production that last figure only falls to 22.8 percent. More than one in five shipped systems have no way to tell whether today’s prompt change made them worse.
We front load this work for a commercial reason. Our $24,500 AI Integration Sprint waives its final milestone if delivery misses day 30, and that guarantee is only enforceable because the eval set is written in week one. A fixed date with no agreed definition of done is a negotiation waiting to happen. The same instinct drives our delivery telemetry, PULSE, which collects more than 400 data points a week per engagement: what gets measured on a schedule gets fixed while it is still cheap. For the wider sequence this fits into, see our guide on how to ship your first production AI feature in 30 days.
Book a Code Review
If you are about to start an AI feature, the useful question is not which eval framework to adopt. It is whether anyone on the team can write down, today, the 20 inputs the feature must get right and the number that counts as passing. A code review works through that with your engineers against the system you already have, and tells you which parts of the plan are specified and which are still hope. Book a Code Review and get the definition of done written before the first sprint, not after it.