Contents
Map

11 · Prompt & Context Engineering

Prompt Evals & Structured Outputs

View as:

Code Lab 01 - Prompt Evals & Structured Outputs

Evaluate prompt variants the way you would evaluate code changes: run zero-shot and few-shot prompts, with and without schema-constrained decoding, over the same labelled test set on a small local model - and let bootstrap confidence intervals tell you which differences are real.

← Back to Overview: Prompt & Context Engineering · Concepts: Core Techniques · Structured Outputs · Prompts in Production

Learning objectives 1.5 hours
By the end of this page you will be able to:
  • Build a labelled eval set and a harness that scores prompt variants on the same items
  • Measure valid-output rate and accuracy with 95% bootstrap confidence intervals
  • Decide whether a prompt change is an improvement using a paired bootstrap on per-item differences
  • Use Outlines to constrain a local model's output to a Pydantic schema
Prerequisites

What's In This Lab

PropertyDetail
ModelQwen/Qwen3-0.6B (instruct, thinking disabled); any Hugging Face chat model via --model
Datatickets.py: 60 support tickets in 4 categories; 8 form the few-shot pool, 52 the test set
Variantszero_free, few_free, zero_constrained, few_constrained
MetricsValid-output rate, accuracy with 95% bootstrap CI, paired bootstrap CI vs the first variant
VerifiedFull run on a laptop CPU/MPS (Apple M-series) in about 2 minutes with torch 2.14, transformers 5.17 and outlines 1.3
Files01-Prompt-Evals-and-Structured-Outputs/{prompt_lab.py, tickets.py, requirements.txt}
flowchart LR
    D["📋 Labelled tickets<br/>(52 test items)"] --> V1["zero_free"]
    D --> V2["few_free"]
    D --> V3["zero_constrained"]
    D --> V4["few_constrained"]
    V1 & V2 & V3 & V4 --> S["📄 Per-item results<br/>valid? correct?"]
    S --> CI["📊 Accuracy ± 95% CI"]
    S --> PB["🔁 Paired bootstrap<br/>vs baseline"]

    style D fill:#e8e2d9,stroke:#ccc4b8
    style S fill:#d8dfe8,stroke:#b0bac8
    style PB fill:#dde4dc,stroke:#b0c4b0

Run It

cd 11-Prompt-Engineering/CodeLabs/01-Prompt-Evals-and-Structured-Outputs
pip install -r requirements.txt

python prompt_lab.py --limit 8          # smoke test
python prompt_lab.py                    # all four variants on 52 items
python prompt_lab.py --model Qwen/Qwen3-1.7B --variants zero_free,few_free

A reference run (Qwen3-0.6B, 52 items):

zero_free          valid=100.0%  acc= 75.0%  95% CI [61.5%, 86.5%]
few_free           valid=100.0%  acc= 78.8%  95% CI [67.3%, 88.5%]
zero_constrained   valid=100.0%  acc= 73.1%  95% CI [59.6%, 84.6%]
few_constrained    valid=100.0%  acc= 78.8%  95% CI [67.3%, 90.4%]
few_free vs zero_free: diff=+3.8%  95% CI [-7.7%, +15.4%]  not significant
zero_constrained vs zero_free: diff=-1.9%  95% CI [-7.7%, +3.8%]  not significant
few_constrained vs zero_free: diff=+3.8%  95% CI [-5.8%, +13.5%]  not significant

Walkthrough - What to Look At

  1. The intervals are wide. With 52 items, a 95% interval on accuracy spans about 25 points. Eyeballing a handful of outputs, or even comparing two averages, can't tell you whether a prompt is better.
  2. Paired beats unpaired. The paired interval on few_free − zero_free is narrower than the gap between the two per-variant intervals suggests, because pairing removes item difficulty. It still spans zero: a +3.8-point gain is not established at this sample size.
  3. Constrained decoding didn't matter here - and that's informative. A one-field JSON object with an enum is easy for a modern instruct model, so free generation was already 100% valid. Constraints earn their keep with nested schemas, smaller or base models, and long outputs (see the first exercise).
  4. Few-shot examples as chat turns. build_messages inserts the examples as prior user/assistant turns rather than one text block - the natural format for chat models.
  5. The label definitions are in the system prompt. Most misclassifications are boundary disagreements (e.g. is "SSO login loops" a bug or account access?). Improving the definitions is often worth more than adding examples - try it and measure.

Check Yourself

Check yourself
0 / 3 answered
  1. Variant A scores 78% and variant B 75% on the same 52 items; the paired 95% CI of A−B is [−5%, +12%]. What do you report?
  2. Why did schema-constrained decoding not improve the valid-output rate in this lab?
  3. Why does the script shuffle the test set with a fixed seed before applying --limit?

Exercises

Exercise - Make constraints matter

Extend TicketLabel with urgent: bool and entities: list[Entity] where Entity has type (product, person or amount) and value. Update the system prompt and the few-shot answers to describe the new fields. Re-run all variants with --max-new-tokens 120 and compare valid-output rates.

Hint

The parser's non-greedy regex \{.*?\} stops at the first closing brace - make it greedy for nested objects.

Solution

In our run on Qwen3-0.6B, zero-shot free generation fell to 88.5% valid (six unparseable outputs), while both constrained variants stayed at 100%. Few-shot free generation was also 100% valid - examples of the exact output shape are a strong format signal too. Constraints guarantee shape at zero cost in examples or tokens; they don't make the category more accurate.

Exercise - Improve the prompt, prove it

Rewrite the category definitions in SYSTEM to address the errors you see (print mismatches to find them). Compare the new prompt against the original with the paired bootstrap. Then add 60 more labelled tickets of your own and repeat. Does the conclusion change?

Hint

Print (ticket, gold, pred) for wrong items in the loop.

Solution

Targeted definition fixes often give a few points; on 52 items the interval will usually still include zero. Doubling the test set narrows the interval by roughly a factor of √2 - the practical lesson is to size eval sets for the effect you need to detect.

Exercise - Scale the model

Run zero_free and few_free with Qwen3-1.7B and Qwen3-4B (GPU or patience required). Plot accuracy against model size with intervals. Is few-shot's benefit larger for smaller models?

Solution

Typically accuracy rises with size and the few-shot gain shrinks, because larger models infer the task from the definitions alone. Report intervals - differences between adjacent sizes may not be significant on 52 items.

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·