Contents
Map

03 · Math for ML

Statistics for Evaluation

View as:

Statistics for Evaluation

Every eval score is an estimate from a finite sample, so it comes with uncertainty - and many published "improvements" are smaller than that uncertainty. This note covers the statistics needed to report and compare LLM evaluations honestly: standard errors and confidence intervals for accuracy, the bootstrap, paired comparisons and McNemar's test, clustered questions, multiple comparisons, how many eval items you need, the unbiased pass@k estimator, and agreement statistics for validating a judge.

Learning objectives 50 min
By the end of this page you will be able to:
  • Compute the standard error and a Wilson confidence interval for an accuracy, and say how it shrinks with more items
  • Run a bootstrap and a paired bootstrap, and apply McNemar's test to two models scored on the same items
  • Explain why clustered items and multiple comparisons make naive intervals too narrow, and what to do about it
  • Estimate the number of eval items needed to detect a given difference
  • Compute pass@k with the unbiased estimator, and Cohen's kappa for judge-human agreement
Prerequisites

An Eval Score Is an Estimate

If a model answers each item correctly with some true probability p, the accuracy on n items is an average of n coin flips. Its standard error is

SE = √( p (1 - p) / n )

Worked example. 160 correct out of 200 (80%): SE = √(0.8 x 0.2 / 200) ≈ 0.028, so the usual 95% interval p ± 1.96 SE is 80% ± 5.5 points. A second model at 83% on the same 200 items is not distinguishable from the first by these intervals alone.

The interval shrinks only with √n: four times the items halves its width.

Use the Wilson interval for small n or extreme p. The normal approximation misbehaves near 0% or 100% and with few items - it can produce intervals below 0 or above 1. The Wilson score interval does not; for 160/200 it gives [73.9%, 85.0%], slightly asymmetric. Bowyer et al. (2025) argue that with fewer than a few hundred items the normal (CLT) interval is unreliable and recommend Wilson-type or Bayesian intervals instead.


The Bootstrap

For metrics without a simple formula - F1, BLEU, a judge's mean score, a median latency - use the bootstrap (Efron, 1979): resample the items with replacement many times, recompute the metric on each resample, and take the 2.5th and 97.5th percentiles of those values as a 95% interval.

import numpy as np

def bootstrap_ci(scores, n_boot=10_000, seed=0):
    """scores: per-item metric values (e.g. 0/1 correctness). Returns the 95% percentile interval of the mean."""
    rng = np.random.default_rng(seed)
    scores = np.asarray(scores, dtype=float)
    idx = rng.integers(0, len(scores), size=(n_boot, len(scores)))
    means = scores[idx].mean(axis=1)
    return np.percentile(means, [2.5, 97.5])

def paired_bootstrap_ci(a, b, n_boot=10_000, seed=0):
    """a, b: per-item scores of two systems on the SAME items. Interval for mean(b - a)."""
    return bootstrap_ci(np.asarray(b, float) - np.asarray(a, float), n_boot, seed)

Paired Comparisons

When two models are scored on the same items, compare them on per-item differences, not by checking whether two separate intervals overlap. Item difficulty is shared, so it cancels in the difference, and the paired interval is usually much narrower. Two tools:

  • Paired bootstrap on b - a per item (above): if the 95% interval of the mean difference excludes 0, the difference is significant at about the 5% level.
  • McNemar's test for pass/fail outcomes: only the items where the models disagree carry information. With b items that only model B gets right and c items that only model A gets right:
χ² = (|b - c| - 1)² / (b + c)        compared with a chi-squared distribution with 1 degree of freedom

Worked example. On 500 items, B alone is right on 30, A alone on 15. χ² = (15 - 1)² / 45 ≈ 4.36, p ≈ 0.037 - a real difference at the 5% level, even though the two accuracies differ by only 3 points.

Building Your Own Evals and the Eval Harness lab apply these to real benchmarks.


Clustered Items

Many eval sets contain groups of related items: several questions about the same document, several test cases for the same function, several turns of the same conversation. Items within a group are correlated, so the effective sample size is smaller than the item count, and the naive standard error is too small.

Fixes: compute clustered standard errors, or bootstrap whole clusters (resample documents, then take all their questions) rather than individual items. Miller (2024) gives the clustered-standard-error formulas for evals and shows they can be much larger than the naive ones on benchmarks built from shared passages.


Multiple Comparisons

Try 20 prompt variants on the same eval set and one will look "significantly" better at the 5% level by chance alone. The same happens when you report the best of many checkpoints, seeds or benchmark subsets.

  • Decide the comparison before looking at the results, and keep a held-out test set you evaluate once.
  • If you must test many hypotheses, correct for it - the Bonferroni correction tests each at α / m for m comparisons; it is conservative but simple.
  • Report how many variants you tried.

How Many Items Do You Need?

To detect a difference Δ between two accuracies around p with 5% significance and 80% power, using independent samples, you need about

n per system ≈ (1.96 + 0.84)² x [p₁(1 - p₁) + p₂(1 - p₂)] / Δ²

Worked example. To detect 80% vs 83% (Δ = 3 points): 7.84 x (0.16 + 0.141) / 0.0009 ≈ 2,600 items per system. Most custom eval sets are far smaller - which is why paired designs matter: scoring both systems on the same items and testing the differences needs far fewer items, because only the disagreements count.

Rule of thumb: a 200-item eval can reliably detect differences of around 10 points; detecting 2-3 points needs thousands of items or a paired design with few disagreements.


pass@k and Non-Determinism

With sampling, a model may solve a problem on some attempts and not others. pass@k is the probability that at least one of k samples is correct. Estimating it by drawing exactly k samples is high-variance; Chen et al. (2021) give an unbiased estimator from n ≥ k samples of which c are correct:

pass@k = 1 - C(n - c, k) / C(n, k)

Worked example. n = 10 samples, c = 3 correct: pass@1 = 1 - C(7,1)/C(10,1) = 0.30; pass@5 = 1 - C(7,5)/C(10,5) = 1 - 21/252 ≈ 0.92.

For agents, the reliability question is the reverse - does it succeed on all k tries? - measured as pass^k (Agent Evaluation & Benchmarks).

Even at temperature 0, outputs can vary between runs (batching and floating-point non-determinism in serving engines). Run important evals more than once and report the spread.


Agreement: Validating a Judge

Before trusting an LLM judge, measure how well it agrees with human labels. Raw agreement overstates it when one label dominates, so use Cohen's kappa, which corrects for chance agreement:

κ = (p_observed - p_chance) / (1 - p_chance)

Worked example. A judge agrees with humans on 85% of items; if both label "pass" half the time at random, chance agreement is 50%, so κ = (0.85 - 0.50) / 0.50 = 0.70 - substantial, but far from the "85%" the raw number suggests. See LLM-as-Judge for the validation workflow.


Check Yourself

Check yourself
0 / 5 answered
  1. A model scores 70% on 100 items. Approximately what is the 95% confidence interval using the normal approximation?
  2. Two models are scored on the same 1,000 questions. Which comparison is most appropriate?
  3. Your eval has 50 documents with 10 questions each. Why is the naive standard error over 500 items too small?
  4. With n = 20 samples per problem and c = 2 correct, what is the unbiased pass@1?
  5. You try 30 system-prompt variants and report the best one as a 2-point improvement with p = 0.04. What's wrong?

Exercises

Exercise - Is the fine-tune better?

Base and fine-tuned models are scored on the same 400 items. Base: 268 correct (67%). Fine-tuned: 284 correct (71%). On 52 items only the fine-tuned model is right; on 36 only the base model is right.

  1. Give 95% normal intervals for each model.
  2. Run McNemar's test.
  3. What do you conclude, and what would you report?
Hint

χ² with 1 degree of freedom: 3.84 is the 5% critical value.

Solution
  1. Base: SE = √(0.67 x 0.33 / 400) ≈ 0.0235 → 67% ± 4.6. Fine-tuned: √(0.71 x 0.29 / 400) ≈ 0.0227 → 71% ± 4.4. The intervals overlap heavily.
  2. χ² = (|52 - 36| - 1)² / (52 + 36) = 225 / 88 ≈ 2.56 < 3.84, so p ≈ 0.11.
  3. Not significant at the 5% level: the 4-point gain is plausible but not established. Report both accuracies with intervals, the paired test and its p-value, and either collect more items or treat the result as inconclusive - don't claim an improvement.
Exercise - Size the eval set

You want to detect whether a prompt change moves accuracy from 90% to 93% with independent samples, 5% significance and 80% power. Roughly how many items per system? What could you change to need fewer?

Solution

7.84 x (0.9 x 0.1 + 0.93 x 0.07) / 0.03² = 7.84 x 0.1551 / 0.0009 ≈ 1,350 items per system. To need fewer: use a paired design (same items, McNemar), which depends only on the disagreements; or focus the eval on the slice where the change is expected to matter, where the effect is larger.

Study Notes

Must-know:

  • SE of an accuracy = √(p(1-p)/n); 95% interval ≈ ± 1.96 SE; width shrinks with √n; 200 items at 80% → ± 5.5 points
  • Use Wilson intervals for small n or extreme p; bootstrap for any metric
  • Compare systems with paired tests on the same items: paired bootstrap, McNemar (only disagreements count)
  • Clustered items need clustered SEs or cluster bootstrap; naive intervals are too narrow
  • Many comparisons manufacture "significant" results: pre-register, correct (Bonferroni), hold out a test set
  • Detecting a 3-point difference independently needs thousands of items; paired designs need far fewer
  • Unbiased pass@k = 1 - C(n-c, k)/C(n, k); agents need pass^k too; rerun evals to see non-determinism
  • Cohen's kappa corrects judge-human agreement for chance

References

Last reviewed: 2026-10

⚡AI-assisted content - always verify, always explore multiple perspectives·