Contents
Map

13 ยท Agent Foundations

Agent Evaluation Basics

View as:

Agent Evaluation Basics

Evaluating an agent means measuring whether it reaches the right end state, reliably, at an acceptable cost - not whether its answer sounds right. This chapter covers the minimum evaluation every agent needs before it ships: outcome checks against environment state, reliability over repeated trials (pass^k), trajectory checks for cost and safety, and a small test set built from real tasks. System-level evaluation, benchmarks and observability come in Production Agents.

Learning objectives 45 min
By the end of this page you will be able to:
  • Distinguish outcome, trajectory and response evaluation, and pick graders for each
  • Compute pass@k and pass^k from repeated trials and explain why agents are judged on pass^k
  • Build a first test set of 30-100 tasks with state-based checks, including tasks the agent should refuse
  • Classify agent failures into actionable categories from traces
Prerequisites

Three Things to Evaluate

flowchart LR
    T["๐Ÿ“‹ Task"] --> A["๐Ÿค– Agent episode"]
    A --> O["๐ŸŽฏ Outcome<br/>Is the environment in<br/>the right end state?"]
    A --> J["๐Ÿงญ Trajectory<br/>Were the steps safe,<br/>efficient, within policy?"]
    A --> R["๐Ÿ’ฌ Response<br/>Is the final message<br/>correct and helpful?"]

    style O fill:#dde4dc,stroke:#b0c4b0
    style J fill:#d8dfe8,stroke:#b0bac8
    style R fill:#e8e0d4,stroke:#c8b89a
WhatQuestionBest graderExample check
OutcomeDid the world end up right?Code: compare environment state to the expected stateOrder O1002 is cancelled and nothing else changed
TrajectoryDid it get there acceptably?Code for hard rules; LLM judge for soft onesNo write before identity lookup; โ‰ค 8 tool calls; no calls to forbidden tools
ResponseDid it tell the user the truth, usefully?Code (regex, exact values) or a validated LLM judgeMentions the refund amount 89.50; doesn't claim actions not taken

Grade outcomes on state, not text. ฯ„-bench compares the database at the end of each conversation with an annotated goal state; SWE-bench runs the repository's tests. Both avoid the commonest agent failure - the reply that claims an action that never happened - which a text-based judge will happily score as correct. Use the response check only for what the state can't show (amounts quoted, explanations given).

Multiple correct trajectories usually exist, so don't require an exact sequence of tool calls. Check invariants instead: required preconditions happened before actions, forbidden actions didn't happen, budgets weren't exceeded.


Reliability: pass@k vs pass^k

Agents are stochastic: the same task can succeed on one run and fail on the next. Run each task n times and count successes c. Two metrics answer different questions:

MetricQuestionEstimator (from n trials, c successes)Use for
pass@kIf I try k times, does at least one succeed?1 โˆ’ C(nโˆ’c, k) / C(n, k)Settings where a verifier can pick the good attempt (code with tests)
pass^k (pass-hat-k)If I run it k times, do all k succeed?C(c, k) / C(n, k)User-facing agents: every customer gets one attempt, and you want every attempt to be right

pass^1 = pass@1 = average success rate. As k grows, pass@k rises and pass^k falls; the gap between them measures inconsistency. Yao et al. (2024) introduced pass^k with ฯ„-bench and found function-calling agents succeeded on under half of retail tasks, with pass^8 below 25% - many tasks were solved sometimes, few always.

A worked example: a task succeeds in 3 of 4 trials.

pass@1 = 3/4              = 0.75
pass^2 = C(3,2)/C(4,2)    = 3/6  = 0.50
pass^4 = C(3,4)/C(4,4)    = 0          (C(3,4) = 0: not all 4 succeeded)
pass@2 = 1 โˆ’ C(1,2)/C(4,2) = 1 โˆ’ 0/6 = 1.00

Averages over a test set hide which tasks are flaky, so always look at the per-task success counts too - the lab prints them.


Building a First Test Set

  1. Start from real tasks. Pull 30-100 examples from logs, tickets or user interviews. Synthetic tasks are fine to fill gaps, but real ones reveal the phrasing and edge cases you didn't imagine.
  2. Cover the categories deliberately:
CategoryExample (retail agent)Why
Happy path, single action"Cancel my rain jacket order"Baseline
Multi-step"Move my pending order and refund the bottle from the other one"Tests completeness
Read-only questions"How much did I spend on delivered orders?"Tests tool use without side effects
Should refuse or declineCancel a shipped order; act on another customer's orderTests policy - the end state must be unchanged
Missing or bad inputUnknown customer emailTests that it stops rather than invents
Ambiguous requests"Cancel my order" when there are twoTests clarification behaviour
  1. Write the expected end state for each task (and, where needed, a response check). Make environments resettable so every trial starts from the same state.
  2. Run n โ‰ฅ 3 trials per task at your production sampling settings; report pass@1 with a confidence interval, pass^k, and per-task counts.
  3. Record cost and latency per episode (model calls, tool calls, tokens, seconds) - a change that raises success by 2 points while doubling tokens may not be worth shipping.

With 50 tasks, a confidence interval on pass@1 is roughly ยฑ10-14 points; treat small differences between variants as noise unless a paired comparison says otherwise (Evaluation & Benchmarks).


Reading Failures

Aggregate scores tell you whether; traces tell you why. Read failing trajectories and tag each with one cause:

Failure classLooks likeUsual fix
Claimed actionReply says "done"; no write tool was calledState-based grading; larger model or higher effort; confirmation generated from tool results
Wrong tool / wrong targetRefunds from the wrong order; calls cancel instead of refundTool descriptions; a lookup step before writes
Policy violationActs on a shipped order or another user's orderEnforce the rule in the tool; state it in instructions
Hallucinated argumentId that appears nowhere in the context"Never guess ids"; validation; list tools
IncompleteDid one of two requested actionsPlan/todo tool; define "done"
Gave up / loopedStopped early, or repeated calls until a boundBetter error messages; repeated-call detection

A failure taxonomy with counts turns evaluation into an engineering plan: fix the most frequent class, re-run, repeat. When one grader is an LLM judge, validate it against 30-50 human-labelled examples before trusting it.


Check Yourself

Check yourself
0 / 4 answered
  1. A task succeeded in 2 of 4 trials. What is its pass^2?
  2. Why is pass^k, not pass@k, the right reliability metric for a customer-service agent?
  3. An LLM judge scores 95% of replies as correct, but a database check finds only 60% of tasks completed. What is the most likely explanation?
  4. Why include tasks the agent should refuse in the test set, and how do you grade them?

Exercises

Exercise - Compute reliability

Five tasks were each run 4 times with successes [4, 4, 3, 1, 0]. Compute pass@1, pass^2, pass^4 and pass@4 averaged over tasks.

Solution

pass@1 = (4+4+3+1+0)/20 = 0.60. pass^2 per task = C(c,2)/6 = [1, 1, 0.5, 0, 0] โ†’ mean 0.50. pass^4 = [1, 1, 0, 0, 0] โ†’ 0.40. pass@4 = [1, 1, 1, 1, 0] โ†’ 0.80. The spread from pass^4 (0.40) to pass@4 (0.80) shows two flaky tasks.

Exercise - Write ten tasks

For an agent you care about, write ten test tasks covering all six categories in this chapter. For each, write the expected end state or response check, and say how you would reset the environment between trials.

Solution

Each task needs an initial state, an instruction, an expected final state (or explicit "unchanged"), and optional response assertions. Resetting is typically a fixture: a fresh in-memory database, a container snapshot, or a sandbox account restored from a template.

Exercise - Build a failure taxonomy

Run the lab with --show-failures, read every failing trace, and tag each with a class from this chapter's table. Which single fix would remove the most failures?

Solution

Count by class. In our reference run (Qwen3-8B, full variant) the failing tasks split into claimed actions (refund_one, refund_shoes, mixed), a policy violation the tools don't enforce (other_customer), a wrong target (already_cancelled) and an unsupported answer (count_items). Claimed actions were the largest class, so the first fix targets them - a harness check or templated confirmations, not prompt tweaks. Then measure the fix: in the lab, a nudge turn fixed mixed but not the two refund tasks, because once the model did call refund_item it picked the customer's pending order - a second, wrong-target error that the claimed action had been hiding. The exercise is the method: taxonomy โ†’ most frequent class โ†’ targeted fix โ†’ re-run.

Study Notes

  • Evaluate outcome (state), trajectory (invariants, budgets) and response (only what state can't show)
  • Grade on environment state - text judges reward claimed actions
  • Check invariants, not exact tool sequences
  • pass@k = at least one of k succeeds: 1 โˆ’ C(nโˆ’c,k)/C(n,k); pass^k = all k succeed: C(c,k)/C(n,k)
  • User-facing agents are judged on pass^k; the pass@k-pass^k gap measures inconsistency
  • Test set: real tasks, all six categories including should-refuse, expected end states, resettable environments, n โ‰ฅ 3 trials, cost and latency recorded
  • Read traces; build a failure taxonomy; fix the biggest class first

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท