Contents
Map

07 · Evaluation & Benchmarks

Q&A Review Bank

View as:

Evaluation & Benchmarks - Q&A Review Bank

21 Q&A pairs. Tags: [Easy] = conceptual recall, [Medium] = design decisions and trade-offs, [Hard] = quantitative reasoning or system design.

Learning objectives 25 min
By the end of this page you will be able to:
  • Answer each question from memory before revealing the answer
  • Explain the reasoning behind each answer - the mechanism or trade-off - not only the fact
  • Identify the chapters you are weakest on and revisit them before the module quiz
Prerequisites
  • The concept notes of this module

Q1 [Easy] Name a benchmark for each of: broad knowledge, hard science reasoning, competition math, real-world coding, and instruction following.

MMLU-Pro (or MMLU), GPQA Diamond, AIME (or MATH-500), SWE-bench Verified, IFEval.


Q2 [Medium] What does it mean for a benchmark to be saturated, and what do you do about it?

Top models cluster near the ceiling or near the label-noise floor, so differences are within noise and the benchmark no longer separates models (MMLU, GSM8K, HumanEval). Compare on harder or fresher benchmarks that still discriminate, and prefer ones with a freshness mechanism.


Q3 [Medium] How do log-likelihood and generative multiple-choice scoring differ?

Log-likelihood scoring picks the option whose text the model assigns the highest probability - no generation or parsing. Generative scoring lets the model produce an answer (often after reasoning) and extracts it with a regex or judge. They measure slightly different things and can differ by several points for the same model.


Q4 [Hard] A model is sampled 10 times and 2 samples pass. Compute unbiased pass@1 and pass@3.

pass@1 = 2/10 = 0.2. pass@3 = 1 - C(8,3)/C(10,3) = 1 - 56/120 ≈ 0.533.


Q5 [Medium] pass@k vs pass^k - when does each matter?

pass@k (at least one of k attempts succeeds) matters when you can verify and pick a success, such as generating code against tests. pass^k (all k attempts succeed, ≈ p^k) measures reliability and matters for agents that must succeed every time without a verifier - an 80% agent has pass^8 ≈ 17%.


Q6 [Medium] How does benchmark contamination happen, and how would you detect it without training-data access?

Benchmark items, answers, solutions or paraphrases get into pretraining crawls, fine-tuning sets or synthetic data. Without data access, compare scores on fresh equivalent items (like GSM1k vs GSM8K) or perturbed versions (rephrased stems, shuffled options), use time-split benchmarks, and try membership-inference signals such as Min-K% Prob.


Q7 [Medium] What are the known caveats of human-preference arenas?

Votes reward style (length, formatting, confidence) as well as correctness; the prompt mix reflects arena users rather than your domain; and selective private testing of many variants can inflate a provider's published ranking. Use arenas to shortlist, not to decide.


Q8 [Medium] List four LLM-judge biases and a mitigation for each.

Position (evaluate both orders, count only consistent wins), verbosity (concision criteria, length-controlled comparisons), self-preference (judge from a different model family or an ensemble), and style/confidence (check correctness as a separate criterion).


Q9 [Medium] How do you validate an LLM judge before trusting it?

Label a sample by hand (ideally two annotators), measure the judge's accuracy and Cohen's κ against the human majority and compare with human-human agreement, inspect the confusion matrix for dangerous error types, probe biases with order swaps and padding, and re-validate whenever the judge, rubric or output distribution changes.


Q10 [Medium] Why decompose a quality rubric into binary criteria?

Narrow, concrete yes/no questions are judged more consistently (by models and humans), are easier to validate, and localize failures; a single 1-10 score blends criteria and drifts.


Q11 [Hard] Your eval has 200 items and your system scores 80%. What is the approximate 95% confidence interval?

SE = sqrt(0.8 × 0.2 / 200) ≈ 0.028; 1.96 × SE ≈ 0.055, so roughly 74.5%-85.5%.


Q12 [Hard] Why use a paired test to compare two systems, and which ones would you use?

Both systems run on the same items, so per-item differences cancel item difficulty; that detects much smaller real differences than comparing two independent intervals. Use a paired bootstrap on per-item score differences, or McNemar's test for pass/fail outcomes; resample by cluster when items share a source document.


Q13 [Medium] How should you build a task-specific eval set?

Sample real inputs stratified by type, add every production failure and bug report, add deliberate edge cases (adversarial, multilingual, should-refuse, very long), write expected behaviour rather than one expected string, and split into a dev set for iteration and a held-out test set for reporting.


Q14 [Medium] Why is needle-in-a-haystack insufficient for long-context evaluation?

Current models mostly pass single-needle retrieval at their advertised lengths, but still fail at multi-needle retrieval, multi-hop tracing and aggregation across the context. RULER and your own multi-part questions at several depths and lengths expose the effective context length.


Q15 [Medium] What belongs in an LLM regression gate in CI?

A fast, representative subset of your eval (including critical slices and past failures) run on every prompt, model or pipeline change; fail the build when a critical slice drops beyond its confidence interval; run the full suite nightly or before release; and track results over time against a pinned baseline.


Q16 [Hard] Two model reports show 71.2 and 73.0 on the same benchmark. List what you would check before believing the 1.8-point gap.

Number of items and resulting confidence interval; shots and prompt; chat template; thinking or effort setting; sampling and aggregation (pass@1 vs majority vote); answer-extraction rules; harness and version; whether a subset was used; tool access; and contamination risk for each model.


Q17 [Easy] What two error types does a safety evaluation measure, and why report both?

Harmful compliance (measured as attack success rate on disallowed requests) and over-refusal (refusals of legitimate but edgy requests, e.g. XSTest-style prompts). Reporting only ASR rewards a model that refuses everything; only over-refusal rewards one that helps with anything.


Q18 [Medium] Why must an ASR number state the attacker budget?

Attack success rises with attacker effort - more attempts per behavior (Best-of-N), more shots (many-shot jailbreaking), more turns (Crescendo), or white-box access (GCG). ASR at 1 attempt and at 100 attempts are different measurements, so compare systems at the same budget, ideally as a curve.


Q19 [Medium] How can the grader distort a red-team result, and how do you guard against it?

A grader that only detects refusals counts any non-refusal as a success, even when the response contains nothing harmful, inflating ASR (shown by StrongREJECT). Grade whether the response actually provides the harmful capability, and validate the grader against human labels like any LLM judge.


Q20 [Hard] An output guard model cuts ASR from 11.5% to 3.0% but raises over-refusal from 2.0% to 6.8%. How do you decide whether to ship it?

Weigh the real traffic mix: attacks are usually rare, so 4.8 points of extra refusals may hurt many more legitimate users than 8.5 points of ASR protects. Look per category - a guard enabled or thresholded only for high-severity categories may keep most of the ASR gain with less refusal. Check the guard's own precision and recall on your traffic and its latency (it may buffer streaming). Decide with the product owner, and record it.


Q21 [Hard] In the Red-Team Harness lab, a hardened system prompt cut ASR from 68% to 29%, but over-refusal rose from 10% to 80%. How would you report this, and what would you try next?

Report both metrics with intervals and the attacker budget, and the paired bootstrap change in ASR (-39 points, CI excluding zero) - then state plainly that the configuration refuses most legitimate questions, so it is not shippable as is. Next: scope the refusal rule narrowly to the protected asset, describe the assistant's normal job first, add examples of questions it should answer, and remove the secret from the prompt entirely (look it up through an authorized tool), then re-measure both metrics on the same attack and benign sets.

⚡AI-assisted content - always verify, always explore multiple perspectives·