Contents
Map

07 ยท Evaluation & Benchmarks

What Benchmarks Measure

View as:

What Benchmarks Measure

Every model card leads with a table of benchmark scores. To read one critically you need to know what each benchmark actually tests, how it is scored, and whether it still separates good models from great ones. This note is a field guide to the benchmarks you will see most, grouped by the capability they target.

Learning objectives 45 min
By the end of this page you will be able to:
  • Name the main benchmark for each capability (knowledge, reasoning, math, code, agentic, instruction following, factuality, long context) and say what it tests
  • Explain how scoring works - multiple choice by log-likelihood vs generation with answer extraction, pass@k, pass^k
  • Recognize benchmark saturation and explain why the leaderboard keeps moving to harder tests
  • Read a model card's benchmark table and ask the right follow-up questions (shots, prompts, thinking budget, harness)
Prerequisites

A Map of Common Benchmarks

mindmap
  root((๐Ÿงช Benchmarks))
    ๐Ÿ“š Knowledge
      MMLU
      MMLU-Pro
    ๐Ÿง  Hard reasoning
      GPQA Diamond
      Humanity's Last Exam
      ARC-AGI
    โž— Math
      GSM8K
      MATH-500
      AIME
    ๐Ÿ’ป Code
      HumanEval
      LiveCodeBench
      SWE-bench Verified
    ๐Ÿค– Agentic
      tau-bench
      Terminal-Bench
      OSWorld
    ๐Ÿ“‹ Instructions and facts
      IFEval
      SimpleQA
    ๐Ÿ“ Long context
      Needle in a haystack
      RULER
BenchmarkCapabilityFormatWhat to know
MMLU (2020)Broad knowledge, 57 subjects4-option multiple choiceSaturated at the frontier; some questions have wrong labels
MMLU-Pro (2024)Knowledge + reasoning10-option multiple choice, harder questionsBuilt to restore headroom over MMLU; less sensitive to prompt wording
GPQA Diamond (2023)Graduate-level science4-option MC, 198 questions written to be "Google-proof"Domain experts score ~65%; skilled non-experts with web access ~34%
Humanity's Last Exam (2025)Frontier-difficulty questions across fields~2,500 questions, MC and short answerDesigned to be hard for current models; also measures calibration
GSM8K (2021)Grade-school math word problemsFree-form, exact numeric matchSaturated; used as a sanity check and for RL demos
MATH / MATH-500 (2021)Competition mathFree-form, answer equivalenceMATH-500 is a representative 500-problem subset
AIME (yearly)Olympiad-qualifier math30 integer-answer problems per yearSmall - scores swing a lot; new years give uncontaminated sets
HumanEval (2021)Function-level PythonUnit tests, pass@kSaturated and widely leaked
LiveCodeBench (2024)Competitive programmingUnit tests; problems tagged by release dateEvaluate only on problems released after the model's cutoff
SWE-bench Verified (2024)Real GitHub issue resolutionAgent edits a repo; hidden tests must pass500 human-validated tasks from the original SWE-bench
ฯ„-bench (2024)Tool-using customer-service agentsSimulated user + APIs; database state checkedIntroduced pass^k (all k attempts must succeed)
IFEval (2023)Following verifiable instructions ("no commas", "3 paragraphs")Programmatic checksCheap and objective; narrow
SimpleQA (2024)Short factual questionsGraded correct / incorrect / not attemptedMeasures hallucination vs abstention
RULER (2024)Long-context retrieval and reasoningSynthetic tasks at increasing lengthsShows effective context is often far below advertised

Agentic benchmarks are covered in depth in Production Agents.


How Scoring Works

Multiple choice: log-likelihood vs generation

  • Log-likelihood scoring (classic harness default for base models): compute the model's probability of each option's text and pick the highest. No generation, no parsing - but it measures something slightly different from how a chat model is used. Variants normalize by length (acc_norm).
  • Generative scoring: the model writes an answer (often after reasoning), and a regex or a judge extracts it. Closer to real use; sensitive to prompt format and to the extraction rule.

The same model can differ by several points between the two - one reason scores from different reports aren't comparable.

pass@k and pass^k

pass@k  = probability that at least one of k samples is correct
          unbiased estimator from n โ‰ฅ k samples with c correct:  1 - C(n-c, k) / C(n, k)
pass^k  = probability that all k independent attempts succeed
          (โ‰ˆ p^k for per-attempt success p) - a reliability measure for agents

pass@k rewards a model that can solve a task; pass^k rewards one that reliably does. An agent with 80% per-attempt success has pass^8 โ‰ˆ 17%.


Saturation and the Moving Frontier

A benchmark stops being useful when top models cluster near its ceiling (or near its label-noise floor): MMLU, GSM8K and HumanEval are there. The community responds with harder or fresher tests - MMLU-Pro, GPQA, HLE, LiveCodeBench, SWE-bench Verified - which then saturate in turn. Two consequences:

  1. Compare models on benchmarks that still separate them. A 0.5-point MMLU difference between frontier models is noise.
  2. Prefer benchmarks with a freshness mechanism (release-dated problems, new yearly sets, private held-out splits) - see Contamination and Leaderboards.

Questions to Ask About Any Score

QuestionWhy it changes the number
How many shots, and which prompt?Few-shot examples and wording can move scores by several points
Thinking on or off, and at what budget/effort?Reasoning models gain most on math/code at high effort
Single sample or majority vote / best-of-N?"maj@64" and pass@1 are different numbers
Which harness and answer extraction?Log-likelihood vs generative, strict vs flexible regex
Full benchmark or a subset?Subsets (e.g. MATH-500) and "verified" splits differ
Tools allowed?Code execution or search changes math and knowledge scores
How many questions?30 AIME problems โ†’ each is 3.3 points; confidence intervals matter

Check Yourself

Check yourself
0 / 3 answered
  1. An agent succeeds on 90% of individual attempts at a task. Roughly what is its pass^5 (all five attempts succeed)?
  2. Why is LiveCodeBench considered more trustworthy than HumanEval for comparing recent models?
  3. Model A reports 88.1 on a benchmark and Model B 87.6. What do you need to know before concluding A is better?

Exercises

Exercise - Compute pass@k

A model is sampled n = 20 times on a coding problem and 3 samples pass. Compute the unbiased pass@1, pass@5 and pass@10 using 1 - C(n-c, k) / C(n, k). Then explain why taking 5 of the 20 samples at random and checking "any pass" gives a noisier estimate.

Solution

pass@1 = 3/20 = 0.15. pass@5 = 1 - C(17,5)/C(20,5) = 1 - 6188/15504 โ‰ˆ 0.601. pass@10 = 1 - C(17,10)/C(20,10) = 1 - 19448/184756 โ‰ˆ 0.895. The combinatorial estimator averages over all subsets of size k; drawing one random subset adds sampling noise.

Exercise - Audit a model card

Take the benchmark table from any recent model release. For each row, record the shots, prompting, thinking/effort setting, aggregation (pass@1 vs majority vote) and harness if stated. Which rows could you reproduce from the information given?

Study Notes

Must-know:

  • Knowledge: MMLU (saturated), MMLU-Pro; hard reasoning: GPQA Diamond, HLE; math: GSM8K (saturated), MATH-500, AIME; code: HumanEval (saturated), LiveCodeBench, SWE-bench Verified; agents: ฯ„-bench, Terminal-Bench, OSWorld; IFEval, SimpleQA, RULER
  • Log-likelihood vs generative scoring can differ by several points for the same model
  • pass@k (any of k succeeds) vs pass^k (all k succeed, โ‰ˆ p^k) - the second measures reliability
  • Always ask: shots, prompt, thinking effort, aggregation, harness, subset, tools, sample size

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท