Contents
Map

11 ยท Prompt & Context Engineering

Core Techniques

View as:

Core Prompting Techniques

Zero-shot instructions, few-shot examples, chain-of-thought, tool-using loops and roles are the working vocabulary of prompting - each was introduced to fix a specific failure, and each has known limits.

Learning objectives 55 min
By the end of this page you will be able to:
  • Decide between zero-shot, few-shot and many-shot prompting for a task, and design a balanced example set
  • Explain why chain-of-thought helps, which task types benefit, and why it is largely redundant with reasoning models
  • Describe the ReAct loop and how native tool calling replaced text-parsed actions
  • Evaluate persona and negative-instruction prompting against the evidence
Prerequisites
flowchart LR
    A["2020<br/>๐Ÿ“„ Few-shot prompting<br/>(GPT-3)"] --> B["2022<br/>๐ŸŽ“ Instruction tuning<br/>makes zero-shot work"]
    B --> C["2022<br/>๐Ÿง  Chain-of-thought"]
    C --> D["2022-23<br/>๐Ÿ”ง ReAct: reason + act"]
    D --> E["2023-24<br/>๐Ÿ› ๏ธ Native tool calling"]
    E --> F["2024-26<br/>๐Ÿ’ญ Reasoning models<br/>think without being asked"]

    style A fill:#e8e2d9,stroke:#ccc4b8
    style C fill:#d8dfe8,stroke:#b0bac8
    style E fill:#dde4dc,stroke:#b0c4b0
    style F fill:#ddd8e4,stroke:#b8b0c8

Zero-Shot Prompting

A task instruction with no examples. It works because instruction tuning (FLAN, InstructGPT and every chat model since) trained models on thousands of tasks phrased as natural-language instructions, installing a general instruction-following skill. On a base model - one that has only been pretrained - zero-shot instructions often fail: the model continues your text rather than obeying it.

Before adding examples, tighten the zero-shot prompt:

  1. Make the label set explicit - "Classify as exactly one of: billing, bug, account_access, feature_request."
  2. Define the labels - most "model errors" on classification are really disagreements about category boundaries.
  3. Specify the output contract - JSON with a schema, enforced by structured outputs where available.
  4. Give the purpose - "used to route tickets to the right team" helps the model resolve ambiguous cases the way you would.

Few-Shot Prompting (In-Context Learning)

Few-shot prompting shows input-output examples before the real input. The model infers the task, format and decision boundary from them with no weight updates - in-context learning.

messages = [
    {"role": "system", "content": "Classify support tickets. Reply with JSON: {\"category\": ...}"},
    {"role": "user", "content": "Ticket: I was charged twice this month."},
    {"role": "assistant", "content": '{"category": "billing"}'},
    {"role": "user", "content": "Ticket: The export button does nothing in Chrome."},
    {"role": "assistant", "content": '{"category": "bug"}'},
    {"role": "user", "content": f"Ticket: {ticket}"},
]

Writing examples as prior user/assistant turns (as above) usually works better with chat models than one long text block.

What the examples actually teach

FindingSourcePractical consequence
Format and label space matter a lot; for smaller models, randomly wrong labels hurt surprisingly littleMin et al. (2022)Get the format and the input distribution right first
Larger models do learn the input-label mapping: with flipped labels they follow the flips and override their priorsWei et al. (2023)With frontier models, wrong or inconsistent labels in examples will be copied - label them carefully
Predictions are biased toward labels that are frequent or recent in the promptZhao et al. (2021)Balance classes; randomise or vary order; evaluate
Example order alone can move accuracy from near-chance to near-bestLu et al. (2022)Never judge a few-shot prompt on one ordering
With long contexts, hundreds or thousands of examples ("many-shot") keep improving results on many tasksAgarwal et al. (2024)The old "over 20 examples, fine-tune instead" rule is outdated - compare many-shot plus prompt caching against fine-tuning on cost and quality

Choosing examples

  • Cover the input distribution, especially the confusable boundary cases - not five easy ones.
  • Keep formatting identical across examples; the model copies inconsistencies.
  • Retrieve examples dynamically for large or varied tasks: embed a pool of labelled examples and insert the nearest neighbours of each input (a retrieval problem - see RAG).
  • Measure. In the code lab, adding eight examples raised accuracy by about 4 points on 52 tickets, with a 95% interval that includes zero - a real effect cannot be distinguished from noise at that sample size.

Chain-of-Thought (CoT)

Chain-of-thought prompting asks for intermediate reasoning before the answer - either by showing worked examples (Wei et al., 2022) or with a trigger such as "Let's think step by step" (zero-shot CoT, Kojima et al., 2022).

Why it works

A transformer does a fixed amount of computation per generated token. A multi-step problem answered in one token must be solved inside a single forward pass. When the model writes intermediate steps, each step becomes context for the next: the generated text acts as working memory, and the total computation grows with the number of tokens written.

flowchart LR
    subgraph NO["Without CoT"]
        Q1["โ“ Problem"] --> A1(["Answer<br/>(one forward pass)"])
    end
    subgraph YES["With CoT"]
        Q2["โ“ Problem"] --> S1["Step 1"] --> S2["Step 2"] --> S3["Step 3"] --> A2(["Answer"])
    end

    style NO fill:#e8e0d4,stroke:#c8b89a
    style YES fill:#dde4dc,stroke:#b0c4b0

When it helps

A meta-analysis of over 100 papers and 14 models (Sprague et al., 2024) found CoT's gains concentrate on math and symbolic or logical reasoning; on most other task types (knowledge recall, commonsense, many classification tasks) the benefit is small. In the original experiments CoT only helped at large scale - small models wrote fluent but illogical chains that made answers worse.

CoT and reasoning models

Reasoning models (see Prompting Reasoning Models and Post-Training: Reasoning Models) are trained with reinforcement learning to produce long internal reasoning before answering. "Think step by step" is redundant for them, and prescribing the steps usually does worse than stating the goal, constraints and output contract. Explicit CoT remains useful for non-reasoning models and when you need the reasoning visible and auditable in the output.


ReAct: Reasoning + Acting

ReAct (Yao et al., 2022) interleaves reasoning with actions: the model writes a Thought, chooses an Action (a tool call), receives an Observation (the tool result), and repeats until it can answer. It is the ancestor of every tool-using agent.

Thought: I need the current population of Tokyo.
Action: search("Tokyo population")
Observation: About 37 million in the metropolitan area.
Thought: I have what I need.
Answer: Roughly 37 million people live in the Tokyo metropolitan area.

The original implementation was text parsing: the model wrote Action: search(...), and application code parsed the line, ran the tool and appended the observation. The model never executes anything - your code does.

Modern APIs replace the parsing with native tool calling: you declare tools with JSON schemas, the model returns structured tool-call blocks (not free text), and you return tool results in a dedicated message type. This removes a whole class of parsing bugs and lets models call several tools in parallel. The loop, its stop conditions and its failure modes are covered in The Agent Loop and Tool Use & Function Calling.


Roles and Personas

"You are an expert tax accountant..." shifts vocabulary, depth, tone and assumed audience. It does not add knowledge the model lacks, and the evidence that it improves correctness is weak: across 162 personas and thousands of factual questions, Zheng et al. (2024) found that adding a persona to the system prompt did not improve accuracy on average, and the best persona for a question was essentially unpredictable.

Use roles for style and audience ("explain to a new engineer who knows Python but not Kubernetes"), and put effort into what actually moves accuracy: clear task definitions, good examples, the right context and evaluation.


Say What You Want, Not Only What You Don't

Positive instructions ("write flowing prose paragraphs") are generally followed more reliably than bare prohibitions ("don't use bullet points"), and model vendors' own prompting guides recommend them. Keep negative instructions for real hard constraints ("never reveal account numbers") and pair them with the positive behaviour you expect instead.


Check Yourself

Check yourself
0 / 4 answered
  1. You give a large frontier model 8 few-shot sentiment examples in which every label is deliberately flipped (positive reviews labelled negative). What does current evidence predict?
  2. For which task is adding 'Let's think step by step' to a non-reasoning model most likely to help?
  3. In a ReAct loop, who executes the tool?
  4. Your few-shot prompt scored 3 points higher than zero-shot on 50 examples. What should you do before shipping it?

Exercises

Exercise - Build a balanced example set

You have 400 labelled tickets across four categories (60% bug, 25% billing, 10% account_access, 5% feature_request). Design an 8-example few-shot set and a procedure to choose it, and say how you would test whether ordering matters.

Solution

Use 2 per class rather than proportional sampling (proportional would give almost no feature_request examples and bias predictions toward "bug"). Within each class pick boundary cases (e.g. "SSO login loops" for account_access vs bug). Test 5 random orderings on a held-out set of at least 100 items and report the spread; if the spread is large, prefer dynamic nearest-neighbour example selection or more examples.

Exercise - Where does CoT pay off?

Using any non-reasoning model, compare direct answering with zero-shot CoT on (a) 30 GSM8K math problems and (b) 30 single-fact trivia questions. Report accuracy and mean output tokens for each. Which task shows the larger gain per extra token?

Solution

Expect a clear gain on GSM8K and little or none on trivia, while CoT multiplies output tokens on both - consistent with Sprague et al. (2024). The cost of CoT is paid on every task; the benefit is not.

Study Notes

Must-know:

  • Zero-shot works because of instruction tuning; tighten label definitions and output contracts before adding examples
  • Few-shot teaches format, label space and mapping; large models copy the labels you show, including wrong ones
  • Label frequency, recency and order all bias outputs - balance, vary order, evaluate
  • Many-shot prompting with long contexts is a real alternative to fine-tuning for some tasks
  • CoT turns generated text into working memory; gains concentrate on math and symbolic reasoning; reasoning models do it internally
  • ReAct = thought/action/observation; native tool calling replaced text-parsed actions
  • Personas change style, not accuracy; prefer positive instructions

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท