Contents
Map

07 · Evaluation & Benchmarks

Appendix - Summary & Key Terms

View as:

Appendix - Evaluation & Benchmarks

What We Learned

  • Benchmark scores depend on task, dataset, harness, prompt and attempt budget.
  • Keep held-out task-specific sets; detect contamination and avoid tuning on the test set.
  • Validate LLM judges against humans and control position, verbosity and other biases.
  • Compare systems on the same items with uncertainty; gate regressions before release.
  • Red-team harmful compliance and excessive refusal, then feed failures into regression tests.

Key Acronyms, Concepts & Jargon

TermShort meaning
Eval / benchmarkA measurement procedure / a standardized task set and scoring protocol.
HarnessCode that runs tasks, prompts models and computes scores.
Contamination / leakageEvaluation content in training / test information influencing development.
LLM-as-judgeA model grades outputs using an explicit rubric.
Rubric / calibrationScoring criteria / checking scores against trusted labels.
CI / paired bootstrapConfidence Interval / resampling matched items to estimate comparison uncertainty.
Precision / recall / F1Correctness among positives / coverage of positives / their harmonic mean.
pass@kProbability of at least one correct solution in k attempts.
Red-teaming / jailbreakAdversarial testing / an attempt to bypass model safeguards.
ASRAttack Success Rate: successful attacks divided by evaluated attacks.
Over-refusalRejecting benign requests that should be answered.
Regression suiteRepeatable tests for behavior that should remain correct.
Prompt injectionUntrusted content attempts to redirect a model's instructions or actions.

Back to section overview

⚡AI-assisted content - always verify, always explore multiple perspectives·