| Eval / benchmark | A measurement procedure / a standardized task set and scoring protocol. |
| Harness | Code that runs tasks, prompts models and computes scores. |
| Contamination / leakage | Evaluation content in training / test information influencing development. |
| LLM-as-judge | A model grades outputs using an explicit rubric. |
| Rubric / calibration | Scoring criteria / checking scores against trusted labels. |
| CI / paired bootstrap | Confidence Interval / resampling matched items to estimate comparison uncertainty. |
| Precision / recall / F1 | Correctness among positives / coverage of positives / their harmonic mean. |
| pass@k | Probability of at least one correct solution in k attempts. |
| Red-teaming / jailbreak | Adversarial testing / an attempt to bypass model safeguards. |
| ASR | Attack Success Rate: successful attacks divided by evaluated attacks. |
| Over-refusal | Rejecting benign requests that should be answered. |
| Regression suite | Repeatable tests for behavior that should remain correct. |
| Prompt injection | Untrusted content attempts to redirect a model's instructions or actions. |