Contents
Map

07 · Evaluation & Benchmarks

LLM-as-Judge

View as:

LLM-as-Judge

Most useful LLM outputs - summaries, answers, plans, code reviews - have no single correct string to compare against. Human grading is accurate but slow and expensive. Using a strong model to grade outputs ("LLM-as-judge") makes evaluation scale, but the judge brings its own biases. This note covers judge designs, the known biases, and how to validate a judge before trusting it.

Learning objectives 40 min
By the end of this page you will be able to:
  • Choose between pointwise (rubric) and pairwise judging, and reference-based vs reference-free grading
  • Name the main judge biases - position, verbosity, self-preference, style - and a mitigation for each
  • Write a judge prompt with an explicit rubric and structured output
  • Validate a judge against human labels and report agreement properly
Prerequisites

Judge Designs

DesignHowGood forWatch out for
Pointwise with a rubricJudge scores one response against explicit criteria (e.g. 1-5 per criterion, or pass/fail per check)Regression tracking, absolute quality barsScore drift; coarse or vague scales
PairwiseJudge picks the better of two responsesComparing two models or promptsPosition bias; doesn't give absolute quality
Reference-basedJudge compares the response with a gold answerFactual QA, extraction, anything with a known answerPenalizes correct answers phrased differently if the rubric is careless
Reference-freeJudge assesses the response on its own meritsOpen-ended writing, helpfulnessJudge's own knowledge limits and preferences

Binary checks beat vague scales. "Does the answer cite the policy section? yes/no" is more reliable than "Rate accuracy 1-10". Decompose quality into a checklist of specific, independently judgeable criteria.

flowchart LR
    IN["📥 Input + response<br/>(+ reference)"] --> J["⚖️ Judge model<br/>rubric · checklist ·<br/>structured JSON output"]
    J --> OUT["📊 Per-criterion verdicts<br/>+ short rationale"]
    OUT --> AGG["🧮 Aggregate<br/>pass rate, win rate"]
    H["👩‍⚖️ Human labels<br/>on a sample"] --> VAL["✅ Validate judge<br/>agreement, κ, bias checks"]
    OUT --> VAL

    style J fill:#d8dfe8,stroke:#b0bac8
    style VAL fill:#dde4dc,stroke:#b0c4b0
    style H fill:#e8e0d4,stroke:#c8b89a

Known Biases and Mitigations

BiasWhat happensMitigation
PositionIn pairwise judging, the judge favours the first (or second) responseEvaluate both orders; count a win only if consistent, otherwise a tie
VerbosityLonger answers score higher regardless of qualityRubric criteria that reward concision; length-controlled comparisons
Self-preferenceA judge rates outputs from its own model family higherUse a judge from a different family, or an ensemble
Style and confidenceConfident, well-formatted answers beat hedged correct onesCriteria that check correctness separately from presentation
Rubric leakageThe model being evaluated was tuned against the same judge promptKeep the evaluation judge separate from any reward model used in training

Zheng et al. (2023) found a strong judge's agreement with human experts on MT-Bench was comparable to agreement between humans, alongside clear position, verbosity and self-enhancement biases. Both findings still hold: judges are useful and measurably biased.


A Judge Prompt Template

You are grading an answer to a customer's billing question.

<question>{question}</question>
<policy_excerpt>{retrieved_policy}</policy_excerpt>
<answer>{answer}</answer>

Evaluate each criterion independently. For each, give a one-sentence reason, then a verdict.
1. correct: every factual claim is supported by the policy excerpt
2. complete: the answer addresses every part of the question
3. actionable: the customer knows exactly what to do next
4. no_invention: the answer does not invent fees, dates, or procedures

Return JSON: {"correct": {"reason": "...", "pass": true|false}, ...}

Put the reasoning before the verdict (so the verdict is conditioned on it), ask for structured output, and keep temperature at 0 (or average several samples) for stability.


Validating a Judge

Before you trust a judge's numbers:

  1. Label a sample by hand - 100-200 items, ideally by two people.
  2. Measure agreement - accuracy against the human majority, and Cohen's κ (chance-corrected); compare with human-human agreement as the ceiling.
  3. Check the confusion matrix - a judge that is lenient on false claims is worse than its accuracy suggests.
  4. Probe biases directly - swap orders, pad answers with filler, and check scores move only when they should.
  5. Re-validate when you change the judge model, the rubric, or the kind of outputs being judged.

Check Yourself

Check yourself
0 / 3 answered
  1. In a pairwise comparison, the judge prefers response A when A is shown first, and B when B is shown first. How should you record it?
  2. Why is 'pass/fail per specific criterion' usually more reliable than a single 1-10 score?
  3. Your judge agrees with human labels 82% of the time. Two human annotators agree with each other 85% of the time. What does that suggest?

Exercises

Exercise - Build and validate a judge

Pick a task you care about (e.g. summarizing support tickets). Collect 60 model outputs and label each pass/fail on three criteria yourself. Write a judge prompt with those criteria, run it, and report per-criterion accuracy and Cohen's κ against your labels. Then run a position-swap or padding probe and report whether the judge's verdicts changed.

Exercise - Reduce verbosity bias

Take 30 pairs where one answer is correct and concise and the other is correct but padded with 3× the text. Measure how often your judge prefers the padded answer, then revise the rubric to fix it and measure again.

Study Notes

Must-know:

  • Pointwise (rubric) vs pairwise; reference-based vs reference-free; prefer decomposed binary criteria
  • Biases: position, verbosity, self-preference, style/confidence, rubric leakage - with a mitigation for each
  • Judge prompt: explicit criteria, reasoning before verdict, structured output, low temperature
  • Validate against human labels (accuracy, κ, confusion matrix) and probe biases before trusting scores

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·