Structured Outputs
Structured outputs make a model's response conform to a schema - JSON, a regex, an enum or a grammar - by constraining which tokens it is allowed to generate, so application code can parse every response.
- Place prompt-only JSON, JSON mode, schema-constrained decoding and strict tool calling on a reliability ladder and pick one for a use case
- Explain how constrained decoding works (token masks from a grammar) and why it can be nearly free at inference time
- Request schema-constrained output from the major hosted APIs and from an open model served with vLLM or Outlines
- Design schemas that don't hurt answer quality (field order, reasoning fields) and add the semantic validation that constraints can't provide
- Prompt Fundamentals
- LLM Fundamentals - decoding and logits
The Reliability Ladder
flowchart LR
A["๐ Prompt only<br/>'Reply in JSON'"] --> B["{ } JSON mode<br/>valid JSON,<br/>any shape"]
B --> C["๐ Schema-constrained<br/>(structured outputs)<br/>valid + matches schema"]
C --> D["โ
+ semantic validation<br/>and retry<br/>valid + schema + correct"]
style A fill:#e8e0d4,stroke:#c8b89a
style B fill:#e8e2d9,stroke:#ccc4b8
style C fill:#d8dfe8,stroke:#b0bac8
style D fill:#dde4dc,stroke:#b0c4b0
| Approach | Guarantees | Typical failures | Use when |
|---|---|---|---|
| Prompt only | Nothing | Preamble text, Markdown fences, trailing commentary, missing or renamed fields | Prototypes; models without constrained decoding |
| JSON mode | Syntactically valid JSON | Wrong keys, wrong types, extra fields | Legacy; prefer schemas |
| Schema-constrained ("structured outputs") | Output parses and matches the schema | Correctly shaped but wrong values; refusals | Any production extraction, classification or data-to-code step |
| Strict tool calling | Tool arguments match the tool's input schema | Wrong tool choice, wrong values | The model is choosing an action, not just returning data |
+ validation and retry (e.g. Pydantic validators, instructor) | Business rules you can check in code | Retries cost latency and money | Values have constraints a JSON schema can't express, or constrained decoding isn't available |
A schema guarantee is about shape, not truth. {"invoice_total": 1200.0} is perfectly valid when the invoice says 120.00. Keep semantic checks (totals add up, dates are in range, IDs exist) in code, and measure accuracy with an eval set.
How Constrained Decoding Works
At every step the model produces logits over the whole vocabulary. A constrained decoder computes which tokens could legally come next given the grammar and what has been generated so far, sets the logits of every other token to minus infinity, and samples from what remains.
flowchart LR
S["๐ JSON schema / regex /<br/>grammar"] --> G["โ๏ธ Compile to automaton<br/>(FSM or pushdown automaton)"]
G --> M["๐ญ Token mask for the<br/>current state"]
L["๐ค Model logits<br/>(whole vocabulary)"] --> X["โ๏ธ Apply mask<br/>illegal tokens โ โโ"]
M --> X
X --> T["๐ฒ Sample next token"] --> U["๐ Advance automaton"] --> M
style G fill:#d8dfe8,stroke:#b0bac8
style X fill:#e8e0d4,stroke:#c8b89a
style T fill:#dde4dc,stroke:#b0c4b0
Key ideas:
- Regular languages (regex, enums, JSON with bounded nesting) compile to a finite-state machine; Willard and Louf (2023, the paper behind Outlines) showed you can precompute, for each FSM state, which vocabulary tokens are allowed, so masking costs a lookup per step.
- Context-free grammars (arbitrary JSON nesting, SQL, code) need a pushdown automaton. XGrammar (Dong et al., 2024) splits tokens into those whose validity is context-independent (precomputed) and a small context-dependent set checked at runtime, bringing overhead to near zero. llguidance takes a similar approach, and both are used by major serving engines.
- Token boundaries are the hard part. Tokens don't align with grammar symbols (one token may close a string, add a comma and open the next key), which is why good libraries work at the token level rather than the character level.
- Constraints can make generation faster: when only one token is legal (a fixed key name, a brace), engines can emit it without sampling.
Hosted APIs
All three major providers support JSON-schema constrained output. The SDKs accept a Pydantic model and return a parsed object.
from pydantic import BaseModel
from typing import Literal
class Ticket(BaseModel):
category: Literal["billing", "bug", "account_access", "feature_request"]
urgent: bool
summary: str
# OpenAI - Responses API
from openai import OpenAI
resp = OpenAI().responses.parse(
model=MODEL, # see the Model Landscape for current names
input=[{"role": "system", "content": "Classify the support ticket."},
{"role": "user", "content": ticket_text}],
text_format=Ticket,
)
ticket = resp.output_parsed
# Anthropic - Messages API
import anthropic
resp = anthropic.Anthropic().messages.parse(
model=MODEL,
max_tokens=1024,
messages=[{"role": "user", "content": f"Classify this support ticket:\n{ticket_text}"}],
output_format=Ticket,
)
ticket = resp.parsed_output
# Google Gemini - Interactions API (google-genai SDK)
from google import genai
interaction = genai.Client().interactions.create(
model=MODEL,
input=f"Classify this support ticket:\n{ticket_text}",
response_format={"type": "text", "mime_type": "application/json",
"schema": Ticket.model_json_schema()},
)
Things to know:
- Schemas are a subset of JSON Schema. Providers typically require every property to be listed in
requiredandadditionalProperties: false, and support only some keywords. Model optional fields as nullable types (str | None) rather than omitting them. - Refusals don't match your schema. If the model declines on safety grounds the response carries a refusal (a
refusalitem or stop reason) instead of your JSON - handle it before parsing. - Truncation breaks the guarantee. If generation hits the output-token limit mid-object, you get invalid JSON. Check the stop reason and size
max_tokensgenerously - for reasoning models the limit also covers thinking. - Structured outputs replace prefill. On models that no longer accept a prefilled assistant turn, schema-constrained output is the supported way to force a format (see Prompt Fundamentals).
- Strict tool calling applies the same guarantee to tool arguments (
strict: trueon the tool definition). See Tool Use & Function Calling.
Open Models
Outlines wraps a Hugging Face model and constrains it to a Pydantic type, regex, enum or grammar (this is what the code lab uses):
import outlines
from transformers import AutoModelForCausalLM, AutoTokenizer
name = "Qwen/Qwen3-0.6B"
model = outlines.from_transformers(AutoModelForCausalLM.from_pretrained(name),
AutoTokenizer.from_pretrained(name))
raw = model(prompt, output_type=Ticket, max_new_tokens=128) # JSON string, guaranteed to parse
ticket = Ticket.model_validate_json(raw)
vLLM exposes the same capability through its OpenAI-compatible server (response_format with a json_schema) and offline through SamplingParams(structured_outputs=StructuredOutputsParams(json=...)), with XGrammar, llguidance or Outlines as the backend (chosen automatically by default). The older guided_json / guided_* parameters were removed in vLLM 0.12 - update code that still uses them.
from vllm import LLM, SamplingParams
from vllm.sampling_params import StructuredOutputsParams
llm = LLM(model="Qwen/Qwen3-8B")
params = SamplingParams(max_tokens=256,
structured_outputs=StructuredOutputsParams(json=Ticket.model_json_schema()))
out = llm.generate([prompt], params)
Do Constraints Hurt Quality?
Tam et al. (2024) reported that strict format constraints can lower accuracy on reasoning tasks compared with free-form answers, especially when the format forces the answer to come before the reasoning. Follow-up analyses (including from the Outlines team) argued most of the loss came from prompt and schema design rather than constraining itself. The practical rules that follow from both sides:
- Order fields so reasoning comes first. Generation is left to right; a schema of
{"reasoning": ..., "answer": ...}lets the model think before committing, while{"answer": ..., "reasoning": ...}forces it to answer first and justify afterwards. - Keep the prompt's instructions consistent with the schema - describe the fields and label definitions in the prompt as well.
- For reasoning models, the thinking happens before the constrained answer, so a minimal answer-only schema is fine.
- For hard reasoning with a non-reasoning model, use two steps: reason freely, then extract into the schema with a cheap call.
- Measure it. In the code lab, constrained and free decoding scored within 2 points of each other on a four-way classification, with overlapping intervals - on your task the answer may differ.
Check Yourself
- JSON mode is enabled, yet your parser rejects some responses with a KeyError. Why?
- How does a constrained decoder prevent invalid output?
- A schema is {"answer": string, "explanation": string} and accuracy on a reasoning task dropped after adding it to a non-reasoning model. What is the cheapest fix to try first?
- Your structured-output call returns text that fails JSON parsing even though constrained decoding is on. Name two likely causes.
Exercises
Extend the code lab's schema to a nested object: {"category": ..., "urgent": bool, "entities": [{"type": "product|person|amount", "value": str}]}. Run free and constrained generation on Qwen3-0.6B over the 52 test tickets. Compare the valid-output rate and accuracy on category.
Hint
Add the new fields to the Pydantic model and update SYSTEM so the prompt describes them.
Solution
With the flat one-field schema, free generation was 100% valid. With the nested schema, zero-shot free generation fell to 88.5% valid in our run while constrained decoding stayed at 100%; few-shot examples of the exact shape also restored validity. The benefit of constraints grows with schema complexity and shrinks with model capability and good examples.
Write a Pydantic model for an invoice (line_items with quantity and unit price, subtotal, tax, total) with validators that check the arithmetic. Wrap an extraction call so that a validation failure triggers one retry with the error message included in the prompt.
Solution
Use a model_validator(mode="after") that checks subtotal equals the sum of quantity ร unit_price (within rounding) and total equals subtotal + tax, raising ValueError with a specific message. On ValidationError, resend the original prompt plus "Your previous output failed validation:
Study Notes
Must-know:
- Prompt-only โ JSON mode (syntax) โ structured outputs (schema) โ strict tools (action arguments); validate semantics in code
- Constrained decoding masks illegal tokens using an automaton compiled from the schema; XGrammar and llguidance make it near free
- Hosted APIs: OpenAI
responses.parse(text_format=...), Anthropicmessages.parse(output_format=...), Geminiresponse_formatwith a schema - vLLM:
response_formatjson_schema orStructuredOutputsParams;guided_*removed in 0.12 - Handle refusals and truncation; schemas support only a subset of JSON Schema
- Put reasoning fields before answer fields; shape โ correctness
References
- Willard & Louf, Efficient Guided Generation for Large Language Models (2023) - the Outlines FSM approach
- Dong et al., XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models (MLSys 2025)
- Microsoft, llguidance
- Tam et al., Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models (EMNLP 2024 Industry)
- OpenAI, Structured Outputs guide
- Anthropic, Structured outputs
- Google, Gemini API structured output
- vLLM, Structured Outputs
- Outlines documentation
Last reviewed: 2026-09