Contents
Map

07 · Evaluation & Benchmarks

Overview

View as:

07 - Evaluation & Benchmarks

Every decision about a model - which one to deploy, whether a fine-tune helped, whether a new checkpoint regressed - is only as good as the evaluation behind it. This module covers how LLMs are measured: public benchmarks and what they actually test, contamination, model-graded evaluation and its biases, long-context evaluation, building an evaluation harness of your own, and safety evaluation and red-teaming.

Learning objectives 6-7 hours
By the end of this module you will be able to:
  • Read a model card's benchmark table critically - what each benchmark measures, and how saturation and contamination distort it
  • Design a task-specific evaluation set and pick metrics for it
  • Use an LLM as a judge while controlling for position, verbosity and self-preference bias
  • Run a standard evaluation harness against a base and a fine-tuned model and interpret the difference
  • Red-team a model or application and report attack success rate and over-refusal with the attacker budget stated
Prerequisites

Chapter Map

#PageTopicLevel
1What Benchmarks MeasureKnowledge, reasoning, math, code, agentic, instruction-following, factuality and long-context benchmarks; scoring; pass@k vs pass^k; saturationBeginner
2Contamination & LeaderboardsHow contamination happens, detection (n-gram, Min-K%, canaries, GSM1k-style variants), time-split evals, arenas and their caveatsIntermediate
3LLM-as-JudgePointwise vs pairwise, rubrics, biases and mitigations, validating a judge against humansIntermediate
4Building Your Own EvalsEval sets from real traffic, metrics per task, confidence intervals and paired tests, long-context evals, CI regression gatesAdvanced
5Safety Evaluation & Red-TeamingHarm taxonomies, attack classes, manual and automated red-teaming (garak, PyRIT, promptfoo), ASR and over-refusal, guard models, safety benchmarksAdvanced
6Q&A Review Bank21 questions across the moduleAll levels

Code Labs

  • Eval Harness - Base vs Tuned - run lm-evaluation-harness on two models and compare them with paired bootstrap confidence intervals.
  • Red-Team Harness - attack a small assistant holding a canary secret with 42 attacks in 7 classes, measure four defence layers (hardened prompt, Qwen3Guard input guard, output filter) by ASR, ASR@k and over-refusal with Wilson intervals and paired bootstrap. Runs on a laptop in minutes.

Agent-level evaluation (trajectories, pass^k, public agent benchmarks) continues in Production Agents.

Section Appendix

Summary & Key Terms - a quick recap of this section and its essential vocabulary.


Next Topic

Previous: 06 - Fine-Tuning Lab · Next: 08 - Inference & Serving

⚡AI-assisted content - always verify, always explore multiple perspectives·