Quiz · 07 · Evaluation & Benchmarks
23 questions from 7 pages
These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
What Benchmarks Measure
Check yourself
0 / 3 answered
- An agent succeeds on 90% of individual attempts at a task. Roughly what is its pass^5 (all five attempts succeed)?
- Why is LiveCodeBench considered more trustworthy than HumanEval for comparing recent models?
- Model A reports 88.1 on a benchmark and Model B 87.6. What do you need to know before concluding A is better?
Contamination & Leaderboards
Check yourself
0 / 3 answered
- A model scores 92% on GSM8K but 78% on GSM1k, a fresh set of matched problems. What is the most likely explanation?
- Which detection method works even when you have no access to the model's training data?
- Why can a model rank highly on a human-preference arena without being more accurate?
LLM-as-Judge
Check yourself
0 / 3 answered
- In a pairwise comparison, the judge prefers response A when A is shown first, and B when B is shown first. How should you record it?
- Why is 'pass/fail per specific criterion' usually more reliable than a single 1-10 score?
- Your judge agrees with human labels 82% of the time. Two human annotators agree with each other 85% of the time. What does that suggest?
Building Your Own Evals
Check yourself
0 / 3 answered
- Your eval has 100 items. System A scores 81%, System B 84%, on the same items. What is the right next step?
- Why keep a separate test split of your own eval set?
- Why is needle-in-a-haystack no longer a sufficient long-context test?
Safety Evaluation & Red-Teaming
Check yourself
0 / 5 answered
- A safety report shows ASR falling from 12% to 1% after a change, and nothing else. What key number is missing?
- Model A has 2% ASR at one attempt per behavior; model B has 6% ASR at 100 attempts per behavior. Which is more robust?
- Why can a refusal-only grader overstate jailbreak success?
- What makes indirect prompt injection different from a jailbreak?
- Your output guard model cuts ASR substantially but users complain that streaming responses now appear all at once after a delay. Why, and what could you do?
Eval Harness - Base vs Tuned
Check yourself
0 / 2 answered
- The base model scores 0.52 (95% CI 0.45-0.59) and the candidate 0.57 (95% CI 0.50-0.64) on the same 200 GSM8K items. The paired difference is +0.05 with 95% CI +0.02 to +0.08. What do you conclude?
- Why might the same model score several points differently in two papers that both report 'GSM8K'?
Red-Team Harness
Check yourself
0 / 4 answered
- Configuration B cuts ASR from 68% to 29%. Why is it not simply 'better' than A?
- The Wilson intervals for B (22-38%) and C (14-28%) overlap, yet the lab concludes the guard has an effect. Why?
- Why does the harness grade with a deterministic leak detector rather than an LLM judge?
- The output filter achieved 0% ASR. What would you need to see before trusting it in production?