Kaikki artikkelit
9. lokakuuta 2026

How to build an eval set for an LLM feature before you ship

Build an LLM eval set before launch: define measurable criteria, mine real cases, add edge cases, validate your grader and gate releases by severity.

Concentric open rings with one acid dot on a dotted canvas
Illustration: Kallos Labs.

Start with success criteria for your LLM evaluation

An eval set is a fixed collection of realistic inputs, each paired with a rule for what a good output looks like, that you run against your feature before every release. Build it before launch, because a feature judged on a handful of demo prompts will fail in ways you only learn about from users.

Start with the criteria, not the cases. Anthropic's guidance on defining success criteria asks for criteria that are specific, measurable, achievable and relevant. "Good performance" is none of those. "Accurate sentiment classification" is closer, and a threshold makes it testable. Anthropic's own example of a measurable safety target: fewer than 0.1% of outputs flagged for toxicity across 10,000 trials.

Most features need several criteria at once. Pick the dimensions that matter for your use case, such as task accuracy, tone, grounding in your sources, latency and cost per request. Write each one as a sentence a stranger could score. If two people would disagree about whether an output passes, the criterion is not finished.

Do this before you compare models, too. The criteria are what turn choosing a model for the feature into a measurement instead of a preference.

Where the eval cases come from

The best cases come from real use. Engineers writing prompts at a desk produce tidy, polite, well-formed inputs. Users send typos, half questions and requests nobody planned for.

Good sources, in rough order of value:

  • Real user questions, support tickets and search logs.
  • Failed prompts and past incidents.
  • Examples from the people who know the domain best.
  • Synthetic cases, used to fill gaps rather than to replace real ones.

OpenAI's evaluation best practices list the same range of sources (synthetic, domain-specific, human-curated, production and historical data) and warn against datasets that "don't faithfully reproduce production traffic patterns." Its simplest instruction is "log everything" while you build, so you can mine those logs for strong cases later. Turn logging on in your prototype, not after launch.

On size, the honest answer is that it depends on risk. One practical guide to pre-ship LLM evaluation suggests a starting golden set of 200 to 500 cases, with fewer for a low-risk summarizer and more for a feature that touches compliance or money. Anthropic's example evals run from 50 groups of paraphrased questions to 1,000 labeled items for a single criterion. If you have 40 labeled cases this week, run those. A small set you trust beats a large one you never finished labeling.

Anthropic also argues for volume over polish: more questions with automated grading, even slightly noisier, beat fewer questions that need hand grading. Automation is what lets you run the set on every change.

Cover the cases that break LLM features

A set made only of happy paths will pass. It will also tell you nothing. OpenAI recommends including typical, edge and adversarial cases. Anthropic names concrete edge cases: irrelevant or nonexistent input, overly long input, poor or harmful input, and ambiguous cases where even humans would struggle to agree.

Checklist diagram of the six items in “Cover the cases that break LLM features”
Checklist: Cover the cases that break LLM features.

Features built on your own documents add their own failure modes. If you are shipping retrieval-augmented assistants, add cases for these:

  • No source: the answer is not in the documents, and the model should say so.
  • Conflicting sources: two documents disagree.
  • Stale source: the document is out of date.
  • Permission boundary: the user is not allowed to see the answer.
  • Tool error: a downstream call fails mid-task.
  • Prompt injection: the input tries to override the instructions.

Write each case as a structured record, not a loose prompt. A workable shape:

  • Input: "Can I get a refund after 45 days?"
  • User role: customer
  • Expected behavior: cites the current refund policy and states the 30-day limit
  • Forbidden behavior: invents an exception, or quotes a retired policy
  • Severity if failed: high
  • Owner: support lead

The role and severity fields matter. The same answer can be right for one user and wrong for another, and a wrong policy answer is not the same size of problem as a missing citation. You will use severity again at the release gate.

Choose graders you can trust

A grader is whatever decides pass or fail for each case. There are three families, and most teams end up combining them.

Code-based graders are deterministic: exact match, string checks, ROUGE-style overlap against reference text, or function-call accuracy for tool use. They are cheap and stable, which makes them good for regression testing. They miss nuance. Anthropic's example eval for task fidelity is a plain exact match on 1,000 human-labeled tweets, and it is a few lines of code.

Human graders are the highest quality and the slowest. Reserve them for building the labeled core of the set and for calibrating everything else. OpenAI suggests showing graders examples of each score level and using a clear pass or fail threshold.

LLM judges scale to cases where nuance matters, such as tone or whether a response used the conversation context. They need care. OpenAI recommends pairwise comparison or pass or fail scoring over open-ended scores, controlling for response length because judges tend to favor longer answers, and having the judge reason before it scores. It also flags position bias and verbosity bias. Anthropic's practical advice is to grade with a different model than the one that generated the output, and to tell the judge to output only a number or "yes" or "no" so the result parses cleanly.

Teams skip one step: validating the judge. Before you trust it, score a labeled sample with the judge and compare against human labels. OpenAI states the rule directly: validate the judge against human labels before scaling. If agreement is poor, fix the judge prompt or switch to a code check. An unvalidated judge is a guess dressed up as a metric.

Also test the eval itself. Feed it outputs you already know are bad and confirm it flags them. An eval that would not have caught your last real incident will not catch the next one.

Gate the release by severity, not by average

Run the set on every change that can alter behavior: a prompt edit, a model upgrade, a retriever change, a new tool schema. Treat each of those as a release and compare results against the previous version. A prompt tweak that improves summaries can quietly worsen refusals, and without a fixed set you will not see it until users do. Nondeterminism makes this more important, not less: OpenAI notes that models can return different outputs for the same input, so a single pass or fail on one run proves little. Score many cases and compare distributions.

Do not gate on one average score. The guide cited above warns that "a single average score can hide serious failures." A feature that passes 96% of cases can still leak data in the other 4%. Set the gate by severity instead:

  • Critical failures (for example, exposing data to the wrong user): zero allowed. Block the release.
  • High-severity failures (a wrong policy answer): fix them, or accept them in writing with a named owner.
  • Medium failures (a missing citation): fix before broad rollout.
  • Low failures (cosmetic): ship if tracked.

For features that take actions, pair the gate with runtime controls. The same severity thinking drives approval gates and audit logs for agents, and the eval set tells you which actions need them.

Keep the set alive after launch

An eval set is never finished. Each production failure should become a new regression case, so the suite gets harder in exactly the places your feature is weak. Keep the set in version control and change it through review, the way you would treat a schema migration. Silent edits make scores look better without making the feature better.

One dated note for tooling choices: as of October 2026, OpenAI's documentation says its hosted Evals platform becomes read-only for existing users on October 31, 2026 and shuts down on November 30, 2026. Keep your cases and graders in files you own, so moving between tools is a configuration change and not a rebuild.

If you are scoping an AI feature and want help designing the evals around it, Kallos Labs covers that work under AI automation.

Conclusion

Write the criteria first, fill the set with real cases and the edge cases that break your feature, pick graders you have checked against human judgment, and block releases on severity instead of an average. Then feed every live failure back in. That loop is the difference between a demo and a feature you can change safely.

Frequently asked questions

How many examples does an eval set need before launch?

It depends on risk, and there is no universal number. One practical guide suggests a starting golden set of 200 to 500 cases, with fewer for low-risk features and more for compliance-heavy ones. Anthropic's example evals use between 50 paraphrase groups and 1,000 labeled items per criterion. Start with what you can label accurately this week, then grow the set from production failures.

Can an LLM grade the outputs of another LLM?

Yes, but validate the judge first. OpenAI recommends checking it against human labels before scaling, controlling for response length because judges favor longer answers, and having the judge reason before it scores. Anthropic advises using a different model to grade than the one that produced the output.

How is an eval set different from unit tests?

Model outputs are nondeterministic, so one pass or fail on one input proves little. An eval set scores many representative inputs against defined criteria and compares results across versions. That comparison is what lets you notice a regression after a prompt or model change.