Contents
Map

Quiz · 07 · Evaluation & Benchmarks

23 questions from 7 pages

These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →

What Benchmarks Measure

Check yourself
0 / 3 answered
  1. An agent succeeds on 90% of individual attempts at a task. Roughly what is its pass^5 (all five attempts succeed)?
  2. Why is LiveCodeBench considered more trustworthy than HumanEval for comparing recent models?
  3. Model A reports 88.1 on a benchmark and Model B 87.6. What do you need to know before concluding A is better?

Contamination & Leaderboards

Check yourself
0 / 3 answered
  1. A model scores 92% on GSM8K but 78% on GSM1k, a fresh set of matched problems. What is the most likely explanation?
  2. Which detection method works even when you have no access to the model's training data?
  3. Why can a model rank highly on a human-preference arena without being more accurate?

LLM-as-Judge

Check yourself
0 / 3 answered
  1. In a pairwise comparison, the judge prefers response A when A is shown first, and B when B is shown first. How should you record it?
  2. Why is 'pass/fail per specific criterion' usually more reliable than a single 1-10 score?
  3. Your judge agrees with human labels 82% of the time. Two human annotators agree with each other 85% of the time. What does that suggest?

Building Your Own Evals

Check yourself
0 / 3 answered
  1. Your eval has 100 items. System A scores 81%, System B 84%, on the same items. What is the right next step?
  2. Why keep a separate test split of your own eval set?
  3. Why is needle-in-a-haystack no longer a sufficient long-context test?

Safety Evaluation & Red-Teaming

Check yourself
0 / 5 answered
  1. A safety report shows ASR falling from 12% to 1% after a change, and nothing else. What key number is missing?
  2. Model A has 2% ASR at one attempt per behavior; model B has 6% ASR at 100 attempts per behavior. Which is more robust?
  3. Why can a refusal-only grader overstate jailbreak success?
  4. What makes indirect prompt injection different from a jailbreak?
  5. Your output guard model cuts ASR substantially but users complain that streaming responses now appear all at once after a delay. Why, and what could you do?

Eval Harness - Base vs Tuned

Check yourself
0 / 2 answered
  1. The base model scores 0.52 (95% CI 0.45-0.59) and the candidate 0.57 (95% CI 0.50-0.64) on the same 200 GSM8K items. The paired difference is +0.05 with 95% CI +0.02 to +0.08. What do you conclude?
  2. Why might the same model score several points differently in two papers that both report 'GSM8K'?

Red-Team Harness

Check yourself
0 / 4 answered
  1. Configuration B cuts ASR from 68% to 29%. Why is it not simply 'better' than A?
  2. The Wilson intervals for B (22-38%) and C (14-28%) overlap, yet the lab concludes the guard has an effect. Why?
  3. Why does the harness grade with a deterministic leak detector rather than an LLM judge?
  4. The output filter achieved 0% ASR. What would you need to see before trusting it in production?