Evals before features.
Most AI teams build the feature, then try to work out whether it works. We build the measurement first. Here’s what that looks like in practice, and why it makes teams faster.
Ask an AI team how they know their system works and you’ll often hear about a demo. Someone typed in twenty questions, the answers looked good, and the feature shipped. Two weeks later a prompt tweak fixes one complaint and quietly breaks three things nobody re-checked.
Evaluation-driven development reverses the order. Before we write application code, we write down what “working” means as a set of examples with expected behavior, and we build the harness that scores the system against them. Every change after that is measured, not admired.
Why the order matters
LLM systems are unusual software. The same input can produce different outputs, small prompt changes have effects far from where you made them, and a model upgrade can shift behavior across thousands of cases at once. Unit tests still matter for the deterministic code around the model, but they can’t tell you whether an answer is right.
Without an eval set, every change is a vibe check. Teams compensate by changing less, which is why so many AI features stall after launch: nobody is confident enough to touch them. With an eval set, a prompt change, a new retrieval strategy or a model swap becomes an experiment with a number at the end.
If you can’t measure a change, you can’t make it safely. And if you can’t make changes safely, you’ve stopped improving the system.
What goes in a first eval set
The first version doesn’t need to be large. It needs to be honest. We usually start with 100 to 300 examples drawn from four places:
- Real traffic. Logs, support tickets, emails, search queries: whatever the system will actually see. Synthetic examples help with coverage, but they’re cleaner than reality.
- The edge cases experts worry about. Ask the people who do the work today what a junior colleague usually gets wrong. Those are your hard cases.
- Things it must refuse. Out-of-scope questions, requests for advice the system isn’t allowed to give, and attempts to make it ignore its instructions.
- Known failures. Every bug report becomes a case. This is the part of the set that grows fastest.
Tag every example with a slice: the kind of input it represents, such as refunds, technical questions or angry customers. Slices matter more than the overall average, because a system can improve on average while getting worse at the one thing your most valuable customers rely on.
Choosing graders
A grader turns an output into a score. Use the cheapest grader that measures what you care about, and save model-based grading for qualities you can’t check any other way.
| Grader | Good for | Watch out for |
|---|---|---|
| Exact or structured match | Classification, extraction, JSON fields, tool choice | Brittle on free text; normalize before comparing |
| Reference checks | Facts that must appear, citations present, numbers within tolerance | Misses answers that are right but phrased differently |
| Rubric graded by a model | Helpfulness, tone, completeness, faithfulness to sources | Needs calibrating against human labels; drifts when the judge model changes |
| Human review | Calibrating the other graders; high-stakes slices | Slow and expensive; use it on samples, not everything |
Model-graded rubrics deserve particular care. Write the rubric as specific yes-or-no questions (“Does the answer state the refund window from the policy document?”) rather than a score out of ten. Have people label a sample, and check how often the judge agrees with them before you trust it. Run that check again whenever you change the judge model.
Scores become release gates
An eval set earns its keep when it can stop a release. We wire it into the deployment pipeline as a gate: every candidate is scored on every slice, and it ships only if it clears an absolute bar and doesn’t regress any critical slice beyond a small tolerance.
from dataclasses import dataclass
@dataclass
class SliceResult:
name: str
score: float # 0..1, mean grader score on this slice
critical: bool # a regression here blocks the release
MIN_OVERALL = 0.90
MAX_REGRESSION = 0.02 # tolerated drop on a critical slice
def release_decision(candidate: list[SliceResult], baseline: dict[str, float]) -> tuple[bool, list[str]]:
"""Ship only if the candidate clears the bar and no critical slice regresses."""
reasons = []
overall = sum(s.score for s in candidate) / len(candidate)
if overall < MIN_OVERALL:
reasons.append(f"overall {overall:.3f} is below {MIN_OVERALL}")
for s in candidate:
before = baseline.get(s.name)
if s.critical and before is not None and s.score < before - MAX_REGRESSION:
reasons.append(f"{s.name} fell from {before:.3f} to {s.score:.3f}")
return not reasons, reasons
The decision comes with reasons, so a blocked release says exactly what got worse. The figure below shows what this looks like across a quarter of release candidates: steady improvement, and one candidate held back because it traded a better average for a worse critical slice.
Keeping the set alive
An eval set is a living asset, not a launch checklist. Three habits keep it useful:
- Every production failure becomes a case. When a user reports a bad answer, the fix isn’t done until the example is in the set and the gate would have caught it.
- Version the data with the code. Scores are only comparable on the same set, so the set lives in the repository and every result records which version it ran against.
- Review disagreements, not just failures. Cases where graders disagree with each other, or with a person, are where your definition of “good” is still fuzzy. Settle them and write the answer down.
Teams sometimes see this as overhead that slows the first release. In our experience it does the opposite. The first version ships a little later, and everything after it ships faster, because changes stop being frightening.
Pick the one workflow that matters most, collect a hundred real examples, and write five yes-or-no rubric questions. That’s enough for a working gate within a week, and a baseline every later change has to beat.