Contents
Map

07 · Evaluation & Benchmarks

Building Your Own Evals

View as:

Building Your Own Evals

Public benchmarks tell you roughly how capable a model is. Whether it is good enough for your product - and whether a new prompt, fine-tune or model version made things better or worse - only your own evaluation can say. This note covers how to build a task-specific eval set, choose metrics, handle statistics honestly, evaluate long context, and wire evals into development as regression gates.

Learning objectives 50 min
By the end of this page you will be able to:
  • Build a task-specific eval set from real traffic, known failures and edge cases, with a train/test discipline
  • Choose metrics per task type - exact match, programmatic checks, reference similarity, judge rubrics
  • Report results with confidence intervals and use paired comparisons to detect real differences
  • Evaluate long-context behaviour beyond needle-in-a-haystack
  • Set up an eval harness and regression gates in CI

The Eval Loop

flowchart LR
    T["📥 Real traffic,<br/>bug reports, edge cases"] --> D["🗂️ Eval set<br/>inputs + expected behaviour<br/>(dev / test split)"]
    D --> R["▶️ Run candidate<br/>(prompt, model, fine-tune)"]
    R --> G["🏅 Grade<br/>code checks · references · judge"]
    G --> S["📊 Report<br/>score ± CI, per-slice,<br/>paired diff vs baseline"]
    S --> DEC{"✅ Ship?"}
    DEC -->|"regression"| FIX["🔧 Fix, add failing<br/>cases to the set"]
    FIX --> D
    DEC -->|"better"| SHIP["🚀 Ship + keep as<br/>new baseline"]

    style D fill:#d8dfe8,stroke:#b0bac8
    style G fill:#e8e0d4,stroke:#c8b89a
    style S fill:#dde4dc,stroke:#b0c4b0

Building the Eval Set

  • Start from reality. Sample real (anonymized) inputs, stratified by type. Add every bug report and production failure as a new case - an eval set is a regression suite that grows.
  • Cover the edges deliberately: empty or huge inputs, adversarial and prompt-injection attempts, multilingual inputs, requests the system should refuse or escalate.
  • Write down expected behaviour, not just an expected string - "must cite the refund policy; must not promise a refund amount".
  • Keep a dev/test split. Iterate prompts against the dev set; report on the test set you did not tune against. Otherwise you are overfitting your own eval.
  • Size it for the decision. 50 cases catch big regressions; detecting a 3-point change needs hundreds (see statistics below).

Choosing Metrics

Task typeBest graderExample
Classification, extraction into fieldsExact match / F1 per fieldTicket routing, invoice fields
Structured outputSchema validation + field checksJSON tool arguments
Math, code, SQLExecution against tests or expected resultsUnit tests pass; query returns expected rows
Grounded QA / RAGJudge with the source document: supported? complete?Policy Q&A
Open-ended writingRubric judge + periodic human reviewSummaries, emails
Multi-step agentsFinal-state checks + trajectory reviewTask completed; forbidden actions not taken

Prefer code-based checks wherever possible; use judges for what code can't check; spot-check both with humans.


Statistics You Can't Skip

An eval score is an estimate. With n items and accuracy p, the standard error is about sqrt(p(1-p)/n):

np = 0.8095% CI (± 1.96 SE)
500.80± 11 points
2000.80± 5.5 points
1,0000.80± 2.5 points
  • Report intervals. "82% (95% CI 77-87%)" is honest; "82%" invites over-reading.
  • Compare paired. When two systems run on the same items, analyse per-item differences (paired bootstrap, or McNemar's test for pass/fail). Pairing cancels item difficulty and detects much smaller real differences than comparing two independent intervals.
  • Cluster when items are related. Several questions about the same document aren't independent; resample by document.
  • Account for sampling noise. At temperature > 0, run several samples per item or fix the seed, and report the variance.

Evaluating Long Context

Needle-in-a-haystack (hide one fact in long filler, ask for it) is now too easy - most models pass it at their advertised length. Better tests:

  • RULER: multiple needles, multi-hop tracing, aggregation and QA at increasing lengths - reveals an "effective" context length often far below the advertised one.
  • Your own long documents: questions whose answers require combining information from several places, placed at different depths.
  • Report by position and length - accuracy at the start, middle and end of the context, at 8K / 32K / 128K.

Tooling

ToolWhat it's for
lm-evaluation-harness (EleutherAI)Standard academic benchmarks on open models; reproducible configs. Used in the Eval Harness lab
Inspect (UK AI Security Institute)Custom evals including agentic, tool-using tasks with sandboxes
promptfooPrompt/model comparison and regression tests in CI
Observability platforms (Langfuse, LangSmith, Braintrust, and others)Datasets, experiment tracking and online evaluation for LLM apps - covered in Production Agents

Regression gates in CI: run the fast subset of the eval on every prompt or model change, fail the build if a critical slice drops beyond its confidence interval, and run the full set nightly or before release.


Check Yourself

Check yourself
0 / 3 answered
  1. Your eval has 100 items. System A scores 81%, System B 84%, on the same items. What is the right next step?
  2. Why keep a separate test split of your own eval set?
  3. Why is needle-in-a-haystack no longer a sufficient long-context test?

Exercises

Exercise - Size an eval

You expect your system to score about 90% and want a 95% confidence interval no wider than ±2 points. How many independent items do you need? What changes if you compare two systems paired on the same items?

Solution

SE must be ≤ 2 / 1.96 ≈ 1.02 points = 0.0102. sqrt(0.9 × 0.1 / n) ≤ 0.0102 → n ≥ 0.09 / 0.000104 ≈ 865 items. For a paired comparison the required size depends on how often the systems disagree - if they disagree on only 10% of items, far fewer items are needed to detect a given difference than two independent 865-item evaluations.

Exercise - Build a regression suite

For an LLM feature you own (or an imagined one), write 30 eval cases: 20 from representative traffic, 5 from known failures, 5 adversarial. For each, write the expected behaviour and choose a grader (code check or judge criterion). Split 20/10 dev/test.

Study Notes

Must-know:

  • Eval sets grow from real traffic, failures and deliberate edge cases; keep a dev/test split
  • Grader per task type: exact match/F1, schema checks, execution, grounded judge, rubric judge, final-state checks
  • SE ≈ sqrt(p(1-p)/n); 200 items at 80% → ±5.5 points; report CIs and use paired comparisons (bootstrap, McNemar)
  • Long context: go beyond needle-in-a-haystack (RULER, multi-part questions by depth)
  • Regression gates in CI: fast subset per change, full suite nightly/pre-release

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·