Contents
Map

17 · Production Agents

Evaluation and Benchmarks

View as:

Agent Evaluation and Benchmarks

Agent Evaluation Basics covered the grading of a single agent on a task set - state-based checks, repeated trials, pass^k. This chapter covers evaluation as a production practice - the suites you keep, where the tasks come from, how graders are validated, how evaluation runs in CI and on live traffic - and the public benchmarks that measure agents, including what their numbers do and don't tell you.

Learning objectives 50 min
By the end of this page you will be able to:
  • Build and maintain capability, regression and safety eval suites from real traffic, with a documented grader for each task
  • Validate an LLM judge against human labels and measure its agreement before trusting it
  • Run offline evals in CI and online evals on production traces, and decide what blocks a release
  • Describe what SWE-bench Verified, τ-bench/τ²-bench, GAIA, WebArena, OSWorld, Terminal-Bench, BFCL and AgentDojo measure, and their limitations
  • Interpret a public benchmark claim critically - harness, attempts, contamination, cost
Prerequisites

The Suites You Keep

SuitePass rate you expectPurposeWhen it runs
RegressionNear 100%Catch anything that used to work and broke - prompt edits, model upgrades, tool changesEvery change (CI), blocking
CapabilityDeliberately lowA hill to climb; shows whether a new model or design really helpsOn demand; before adopting changes
Safety and security100% on must-refuse and injection casesPolicy violations, prompt injection, data leakage (e.g. AgentDojo-style tasks)Every change, blocking
OnlineTracked as a trendSampled live traces scored by graders and user signalsContinuously

Tasks move between suites: a capability task the agent now passes reliably becomes a regression task; a production failure becomes a new task in whichever suite fits. Suites saturate - when one sits at 100% it no longer tells you anything new, so add harder cases.

Where Tasks Come From

  1. Real traffic first. Sample production conversations, including the failures users reported and the runs that hit bounds or needed escalation. Real inputs have the phrasing, ambiguity and edge cases synthetic ones lack.
  2. Specified behaviours. For each policy rule, at least one task that tests it, including should-refuse cases (Lab 13's cancel_shipped, other_customer).
  3. Synthetic expansion - a model generates variations of real tasks (paraphrases, different customers, harder combinations) - reviewed by a person before it enters a suite.
  4. Adversarial cases from red-teaming: injected content in tool results, conflicting instructions, social engineering.

Each task needs an outcome check you can defend: final environment state where possible (database rows, files, test results), a reference answer with a tolerant comparison, or a rubric. Write the reason each check exists next to it; graders drift out of date as the product changes.

Graders You Can Trust

State checks and tests are the gold standard: deterministic and cheap. Where you need an LLM judge (answer quality, tone, groundedness, whether an explanation to the customer was adequate):

  • Use a specific rubric, scored per criterion ("mentions the refund amount", "does not promise a date"), not a 1-10 overall score.
  • Validate against humans: have people label 50-200 outputs, run the judge on the same outputs, and measure agreement (accuracy, Cohen's kappa, or correlation). Improve the rubric until agreement is close to agreement between humans.
  • Watch the known biases: position, verbosity and self-preference (judges favour outputs from their own model family) - see Module 07. Use a different model family as judge where possible.
  • Re-validate after changing the judge model or the rubric.

Anthropic's multi-agent research system used an LLM judge with a rubric (factual accuracy, citation accuracy, completeness, source quality, tool efficiency) and found a single judge call scoring 0-1 plus pass/fail was most consistent with human judgement - while humans still caught failure modes the automated evals missed, such as a preference for SEO-farm sources.

Offline and Online

flowchart LR
    CH["✏️ Change<br/>prompt, model, tool"] --> OFF["🧪 Offline evals in CI<br/>regression + safety (blocking)<br/>capability (report)"]
    OFF -->|"pass"| CAN["🐤 Canary / shadow<br/>small % of traffic"]
    CAN --> ON["📈 Online evals<br/>sampled traces scored,<br/>user feedback, escalations"]
    ON -->|"failures"| DS["🗂️ Add to eval suites"]
    DS --> OFF

    style OFF fill:#d8dfe8,stroke:#b0bac8
    style ON fill:#dde4dc,stroke:#b0c4b0
    style DS fill:#e8e0d4,stroke:#c8b89a
  • Offline: run suites on every change with several trials per task; report success with confidence intervals, pass^k, cost and latency per task; block on regressions outside the noise band.
  • Canary or shadow: new versions take a small share of traffic (or run in parallel without acting) before full rollout; compare the metrics.
  • Online: score sampled production traces with the same graders (those that don't need ground truth), and track user signals - corrections, escalations, abandonment, thumbs-down. Alert on trends.
  • Close the loop: every production failure you understand becomes an eval task.

Compare variants on the same tasks and trials (paired comparisons), as the Module 15 lab does: paired confidence intervals are much tighter than comparing two independent pass rates.

Public Agent Benchmarks

BenchmarkWhat it measuresHow it's gradedWatch out for
SWE-bench Verified (2024)Resolving real GitHub issues in Python repositories; 500 tasks human-validated as solvableHidden unit tests pass after the agent's patchPython-only, popular repositories (contamination risk); scores depend heavily on the harness
τ-bench / τ²-bench (2024, 2025)Customer-service agents (retail, airline, telecom) with a simulated user and policiesFinal database state; pass^k over repeated trialsτ-bench showed even strong agents below 50% task success, with pass^8 below 25% in retail; τ² adds dual control, where the user also acts
GAIA (2023)General assistant questions needing browsing, files and reasoning; 466 questionsExact-match answersAt release humans scored 92% vs 15% for GPT-4 with plugins; now largely saturated at the easier levels
WebArena (2023)Tasks on realistic self-hosted websites (shopping, forums, GitLab, maps)Functional checks of site stateAt release the best GPT-4 agent reached 14.4% vs 78.2% for humans
OSWorld (2024)Computer use across real desktop applicationsExecution-based checks per taskAt release humans reached 72.4% vs 12.2% for the best model
Terminal-Bench (2025)Tasks in a real terminal: building, debugging, system administrationTests in the containerMeasures the agent harness as much as the model
BFCL (Berkeley Function-Calling Leaderboard)Function-calling accuracy: single, parallel, multi-turn, irrelevance detectionAST and execution checksTool-call correctness, not end-to-end task success
AgentDojo (2024)Utility and security under prompt injection across email, banking, travel and workspace tasksTask and attack successA defence that lowers attack success but also utility may not be a win
METR time horizons (2025)The length of tasks (in human time) agents complete with 50% reliabilityTask success vs human completion timeReported doubling roughly every 7 months since 2019; wide error bars, software tasks

Reading a benchmark claim:

  • Which harness and scaffold? The same model can differ by many points across harnesses; leaderboards increasingly report the agent system, not the model.
  • How many attempts? pass@1 with one run, best-of-n, or averaged over trials? Is test-time compute (parallel attempts plus selection) included?
  • Contamination: public repositories and web pages leak into training data; prefer benchmarks with held-out or refreshed tasks.
  • Cost and time per task: a score that needs 10× the tokens may not be the better system for you.
  • Relevance: none of these is your task. Use them to shortlist models; decide with your own suite.

Check Yourself

Check yourself
0 / 4 answered
  1. A regression suite passes at 97% after a prompt change, down from 100%. The change improved the capability suite. What do you do?
  2. How do you know an LLM judge is good enough to use?
  3. Why are paired comparisons preferred when comparing two agent versions?
  4. Model A scores higher than model B on SWE-bench Verified. Name three reasons this may not predict which is better for your coding agent.

Exercises

Exercise - Build the suites

For the Lab 13 shop agent, build a regression suite (tasks it passes 4/4), a capability suite (tasks it fails) and a safety suite (should-refuse and injection tasks from Lab 14). Wire them into a script that exits non-zero when regression or safety pass rates fall outside their noise band.

Solution

Use the lab's per-task pass^4 to split tasks; add Lab 14's poisoned-result variant as safety tasks with the attack-success check. The gate: regression pass rate below its historical mean minus the bootstrap CI half-width, or any safety failure, fails the build. Capability results are reported but don't block.

Exercise - Validate a judge

Write a rubric judge for the shop agent's replies (states the outcome accurately, gives the refund amount when relevant, doesn't promise what the tools didn't do). Label 40 replies from Lab 13 yourself and measure the judge's agreement with your labels. Where does it disagree?

Solution

Expect good agreement on "gives the amount" and poor agreement on "doesn't promise what the tools didn't do" unless the judge sees the trace - a reply-only judge can't detect claimed actions. Give the judge the tool results, or replace that criterion with a state check.

Study Notes

  • Suites: regression (blocking), capability (hill to climb), safety (blocking), online (trend); tasks move between them; retire saturated ones
  • Tasks from real traffic, specified behaviours, reviewed synthetic variations, red-teaming; every task has a defended outcome check
  • LLM judges: specific rubrics, validated against human labels, re-validated on change, aware of biases
  • Offline CI -> canary -> online scoring -> failures back into suites; paired comparisons with CIs
  • Benchmarks: SWE-bench Verified, τ/τ²-bench (pass^k), GAIA, WebArena, OSWorld, Terminal-Bench, BFCL, AgentDojo, METR time horizons; check harness, attempts, contamination, cost, relevance

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·