Contents
Map

07 Β· Evaluation & Benchmarks

Contamination & Leaderboards

View as:

Contamination and Leaderboards

A benchmark score means "the model can do this kind of task" only if the model hasn't seen the test. Web-scale pretraining makes that hard to guarantee: benchmark questions, answers and discussions end up in Common Crawl. This note covers how contamination happens, how to detect and avoid it, and how to read public leaderboards - including preference arenas - without being misled.

Learning objectives 40 min
By the end of this page you will be able to:
  • Explain how benchmark contamination happens and why it inflates scores
  • Describe detection methods - n-gram overlap, perplexity/membership tests, canary strings, performance on fresh variants
  • Design evaluations that resist contamination (time-split, private held-out, perturbed variants)
  • Read preference-arena leaderboards critically, including style effects and selective reporting
Prerequisites

How Contamination Happens

flowchart LR
    B["πŸ“‹ Benchmark published<br/>questions + answers"] --> W["🌐 Copied to the web<br/>GitHub, blogs, forums, solutions"]
    W --> CC["πŸ•ΈοΈ Crawled into<br/>pretraining data"]
    B --> FT["πŸŽ“ Used (knowingly or not)<br/>in fine-tuning / synthetic data"]
    CC --> M["πŸ€– Model memorizes<br/>questions or answers"]
    FT --> M
    M --> S["πŸ“ˆ Inflated benchmark score"]

    style M fill:#ddd8e4,stroke:#b8b0c8
    style S fill:#e8e0d4,stroke:#c8b89a

Contamination ranges from verbatim (the exact question and answer in training data) to indirect (paraphrases, translated copies, solution write-ups, or synthetic data generated by a model that itself saw the test). Indirect contamination is harder to detect and increasingly common.

Evidence it matters: when researchers wrote GSM1k - new grade-school math problems matched in style and difficulty to GSM8K - several model families scored markedly lower on it than on GSM8K, consistent with overfitting to the public set (Zhang et al., 2024).


Detecting Contamination

MethodHow it worksLimits
N-gram overlapSearch the training corpus for long n-gram matches with test items (e.g. 13-grams)Needs corpus access; misses paraphrases
Membership inference (e.g. Min-K% Prob)Test items the model saw tend to have unusually few very-low-probability tokensNoisy; works best for verbatim leaks
Canary stringsBenchmarks embed a unique GUID; if the model can complete it, the data leaked (BIG-bench does this)Only works if the canary was kept intact
Fresh or perturbed variantsCompare scores on the public set with new, equivalent questions (GSM1k), or with rephrased / reordered-option versionsBuilding equivalent sets is expensive
Time-splitEvaluate only on items created after the model's training cutoff (LiveCodeBench, new AIME years)Requires ongoing benchmark maintenance

Designing Contamination-Resistant Evals

  1. Keep a private held-out set that never touches the internet or any training pipeline - including synthetic-data generators.
  2. Prefer time-split benchmarks and report which window you used.
  3. Perturb - rephrase questions, shuffle options, change numbers - and treat a large drop as a red flag.
  4. Decontaminate training data against every benchmark you plan to report (see the data pipeline's decontamination stage).
  5. Report the gap between public and private scores; a big gap is itself a finding.

Leaderboards

Static leaderboards

Aggregators such as Artificial Analysis run a fixed suite under a consistent harness and report speed and price alongside quality - useful because they remove the "every vendor used a different setup" problem. They still inherit each benchmark's saturation and contamination issues.

Preference arenas

Chatbot Arena / LMArena collects pairwise human votes on anonymous model responses and fits Bradley-Terry ratings. It measures what real users prefer in open-ended chat, which static benchmarks miss. Known caveats:

  • Style effects. Longer, more formatted, more confident answers win votes independent of correctness; LMArena introduced style control to adjust for length and markdown.
  • Prompt distribution. Votes reflect the kinds of prompts arena users submit - weighted toward chat, coding help and creative writing, light on specialist domains.
  • Selective reporting. "The Leaderboard Illusion" (Singh et al., 2025) documented providers privately testing many model variants and publishing only the best, and unequal access to arena data - both inflate rankings.

How to use leaderboards well: as a coarse filter for a shortlist, never as the final decision. The deciding evaluation is the one you build on your own task (see Building Your Own Evals).


Check Yourself

Check yourself
0 / 3 answered
  1. A model scores 92% on GSM8K but 78% on GSM1k, a fresh set of matched problems. What is the most likely explanation?
  2. Which detection method works even when you have no access to the model's training data?
  3. Why can a model rank highly on a human-preference arena without being more accurate?

Exercises

Exercise - Perturbation test

Take 50 questions from a public multiple-choice benchmark. Create a perturbed version of each (shuffle the options, rephrase the stem without changing its meaning). Evaluate one open model on both versions. Report the accuracy gap with a bootstrap confidence interval, and interpret it.

Exercise - Build a canary

Your team is publishing an internal benchmark to a partner. Write a data-handling policy (half a page) covering a canary string, a private held-out split, a time-split refresh cadence, and how you will detect leakage in future models.

Study Notes

Must-know:

  • Contamination: benchmark text (or paraphrases, solutions, synthetic copies) in training data inflates scores
  • Detection: n-gram overlap, membership inference (Min-K% Prob), canary strings, fresh/perturbed variants (GSM1k), time-split
  • Defences: private held-out sets, time-split benchmarks, perturbation tests, decontaminating training data, reporting public-vs-private gaps
  • Arenas measure human preference in chat, with style and selection effects; use leaderboards to shortlist, not to decide

References

Last reviewed: 2026-09

⚑AI-assisted content - always verify, always explore multiple perspectives·