Contents
Map

11 · Prompt & Context Engineering

Prompting Reasoning Models

View as:

Prompting Reasoning Models

Reasoning models generate a hidden or summarised chain of reasoning before their answer, so prompting them shifts from guiding the reasoning to specifying the goal, the constraints and the output - and to choosing how much reasoning to pay for.

Learning objectives 40 min
By the end of this page you will be able to:
  • Set the reasoning-effort controls of the major APIs and explain what they trade off
  • Rewrite a prompt written for a non-reasoning model into one suited to a reasoning model
  • Handle reasoning state correctly in multi-turn and tool-calling loops
  • Explain why more reasoning is not always better and why visible reasoning is not a faithful explanation
Prerequisites

What Changed

Reasoning models are trained with reinforcement learning to think at length before answering, and their accuracy on hard problems rises with the amount of reasoning they do (test-time compute). By 2026 most frontier models reason by default, and the old distinction between "chat models" and "reasoning models" has become a setting: how much to reason.

Prompting habit from non-reasoning modelsWith reasoning models
"Think step by step"Redundant - the model already reasons
Few-shot examples of reasoningOften worse than stating the goal; examples of the output format still help
Prescribing the procedure ("first list X, then check Y...")Usually worse - give the goal, constraints and success criteria and let the model plan
Tuning temperatureOften unavailable; tune reasoning effort instead
"CRITICAL: YOU MUST..." in capitalsNewer models follow instructions closely and can over-apply shouted rules; state them plainly with the reason
Assistant prefill to force formatRemoved on several APIs; use structured outputs

What did not change: clear task definitions, the right context, explicit output contracts and evaluation matter as much as ever.


The Effort Dial

Each provider exposes a control for how much the model reasons. Higher settings improve accuracy on hard tasks and cost more output tokens and latency.

ProviderControlValues (model-dependent)
OpenAIreasoning={"effort": ...} (Responses API), reasoning_effort (Chat Completions)none, minimal, low, medium, high, xhigh, max
AnthropicAdaptive thinking (thinking={"type": "adaptive"}, on by default on newer models) plus an effort settinglow, medium, high, xhigh, max
Google Geminithinking_level in the generation configlow, medium, high

Older interfaces used a token budget instead of a level (Anthropic's budget_tokens, Gemini 2.5's thinking_budget); current models favour levels, because the model is better than a fixed number at judging how hard a given request is.

# OpenAI Responses API
resp = client.responses.create(model=MODEL, reasoning={"effort": "low"}, input=prompt)

# Anthropic Messages API (adaptive thinking is the default on current models)
resp = client.messages.create(model=MODEL, max_tokens=16000,
                              output_config={"effort": "low"},
                              messages=[{"role": "user", "content": prompt}])

Treat effort as a hyperparameter to evaluate, per task: run your eval set at two or three levels and plot accuracy against cost per item. Classification and extraction are often as accurate at the lowest level; planning, math, debugging and long agentic tasks benefit from more.

Practical notes:

  • Reasoning tokens are billed as output tokens and count against the output limit. Set max_tokens high enough for thinking plus the answer, or you will get truncated outputs.
  • Latency: time-to-first-visible-token includes the thinking. Stream, and show progress (reasoning summaries) in interactive products.
  • Model-specific rules: some models reject certain combinations (for example, disabling thinking at the highest effort levels). Check the provider's model page.

Writing the Prompt

Before (written for a non-reasoning model)
  You are an expert analyst. Think step by step.
  Step 1: List every metric in the report.
  Step 2: For each metric compare to last quarter.
  Step 3: Decide which changes matter.
  Step 4: Write the summary.

After (for a reasoning model)
  Goal: a 5-bullet summary of what changed this quarter, for the CFO.
  What matters: changes over 5% in revenue, gross margin, churn or cash
  burn, and anything the report itself flags as a risk.
  Constraints: use only figures in the report; say "not reported" rather
  than estimating.
  Output: JSON {"bullets": [...], "risks": [...]}.
  <report>...</report>

Guidelines:

  1. State the goal and why it matters - the model resolves ambiguity better when it knows the purpose.
  2. Give success criteria and constraints, not a procedure.
  3. Provide all the context the task needs (documents, definitions, prior decisions). Reasoning cannot recover missing facts - it can only hallucinate them more elaborately.
  4. Specify the output contract and enforce it with structured outputs.
  5. Keep examples for format and style, not for reasoning.
  6. For agentic tasks, say when to stop and what "done" means, and whether to ask or proceed when information is missing.

Reasoning State in Multi-Turn and Tool Loops

When a reasoning model calls a tool, its reasoning so far is part of the state it needs to continue. Each API has a mechanism for carrying it across turns, and dropping it degrades quality or raises errors:

ProviderWhat to carry forward
OpenAIReasoning items: reference them with previous_response_id, or pass them back (including their encrypted_content in stateless mode) with the tool results
AnthropicPass the assistant turn's thinking blocks back unchanged along with the tool_use blocks and your tool_results; on some newer models, editing earlier turns invalidates those blocks
GeminiThought signatures, which the SDK carries automatically in stateful mode

Most APIs return a summary of the reasoning or an encrypted form, not the raw tokens. Don't build features that depend on reading raw reasoning.


More Thinking Is Not Always Better

  • Inverse scaling. Gema et al. (2025) built tasks on which accuracy falls as reasoning length grows - the model gets distracted by irrelevant details, overfits to familiar framings or drifts. Measure accuracy at several effort levels rather than assuming the highest is best.
  • Overthinking simple tasks wastes tokens and latency; use low effort or a non-reasoning setting for routing, classification and extraction.
  • Visible reasoning is not a faithful explanation. Chen et al. (2025) found reasoning models often fail to mention hints that demonstrably changed their answers. Use reasoning summaries for debugging and user feedback, not as an audit trail of why the model decided something - log inputs, tool calls and outputs instead.

Check Yourself

Check yourself
0 / 4 answered
  1. You migrate a classification prompt from a non-reasoning model to a reasoning model. Which change is most likely to help?
  2. A reasoning model's answers are cut off mid-JSON although the answer itself is short. What is the likely cause?
  3. In a tool-calling loop you strip the thinking blocks from the assistant message before sending the tool results back. What happens?
  4. Why shouldn't you treat a model's reasoning summary as the explanation for its decision in an audit?

Exercises

Exercise - Effort sweep

Pick a reasoning model you can access and two tasks: 50 support-ticket classifications and 30 multi-step math problems. Run each at three effort levels. Record accuracy, mean output tokens (including reasoning) and p50 latency. Recommend a level per task.

Solution

Expect classification to be flat across levels (choose the lowest) and math to improve with effort, with diminishing returns at the top - and output tokens rising several-fold. The recommendation should be per task, justified by accuracy per dollar, which is the point of treating effort as a hyperparameter.

Exercise - Rewrite for a reasoning model

Take a long procedural prompt from one of your own projects (or the "before" example above) and rewrite it as goal, context, constraints, success criteria and output contract. Compare the two versions on at least 30 eval items.

Solution

The rewrite is usually shorter and at least as accurate on a reasoning model. If the procedural version wins, look for domain knowledge hidden in the steps (e.g. "check the footnotes for restatements") and move it into the context or constraints rather than restoring the procedure.

Study Notes

Must-know:

  • Reasoning models think before answering; prompts specify goal, context, constraints and output rather than steps
  • Effort controls: OpenAI reasoning.effort, Anthropic adaptive thinking + effort, Gemini thinking_level; token budgets are the older style
  • Reasoning tokens are billed as output and count toward max_tokens
  • Carry reasoning state (reasoning items, thinking blocks, thought signatures) through tool loops
  • More reasoning can hurt (inverse scaling); tune effort per task on an eval set
  • Reasoning summaries are not faithful explanations

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·