Contents
Map

15 ยท Agent Patterns & Multi-Agent

Single-Agent Patterns

View as:

Single-Agent Patterns

Before reaching for several agents, most problems are better served by one well-equipped agent plus a few structural patterns around it: guardrails running alongside, gates before risky actions, sub-agents used as tools for context isolation, handoffs to humans, and verification loops driven by real feedback. This chapter collects those patterns and states when each is worth its cost.

Learning objectives 45 min
By the end of this page you will be able to:
  • Explain why a single agent with good tools is the default architecture and what signals justify more
  • Apply guardrail, gating and human-in-the-loop patterns around an agent loop
  • Use sub-agents as tools for context isolation, and know how they differ from multi-agent collaboration
  • Design verification loops that use external feedback, and escalation paths that hand off cleanly to humans
Prerequisites

The Default: One Agent, Good Tools

A single agent loop with a well-designed tool set, a clear system prompt and a capable model is the right starting point for most tasks. It has one context (no information lost between agents), one trace to debug, and the lowest token cost. Planning strategies (ReAct, plan-and-execute, todo lists) are variations within this agent (Planning & Reasoning).

Move beyond it only when you observe a specific limit:

SymptomLikely remedy
Context fills with tool output the agent no longer needsSub-agents as tools; compaction (Agent Memory)
Too many tools, wrong ones chosenTool search, routing to specialised agents
Independent sub-tasks are slow in sequenceParallel tool calls, or parallel sub-agents
Quality varies and can be checkedA verification loop with external feedback
Some actions are too risky to automateGates and human approval

Guardrails Around the Loop

Guardrails are checks outside the model, running before, during or after it:

flowchart LR
    U["๐Ÿ‘ค Input"] --> IG{"๐Ÿ›ก๏ธ Input guardrails<br/>scope, PII, injection heuristics"}
    IG -->|"blocked"| X["Refuse / escalate"]
    IG -->|"ok"| A["๐Ÿค– Agent loop"]
    A --> TG{"๐Ÿšง Action gate<br/>per write tool"}
    TG -->|"allowed"| T["๐Ÿ”ง Tool"]
    TG -->|"needs approval"| H["๐Ÿ™‹ Human"]
    T --> A
    A --> OG{"๐Ÿ›ก๏ธ Output guardrails<br/>policy, grounding, format"}
    OG -->|"ok"| R["โœ… Reply"]
    OG -->|"fail"| A

    style IG fill:#e8e0d4,stroke:#c8b89a
    style TG fill:#e8e0d4,stroke:#c8b89a
    style OG fill:#e8e0d4,stroke:#c8b89a
    style A fill:#d8dfe8,stroke:#b0bac8
  • Run cheap checks in parallel with the main call (a small classifier for off-topic or abusive input) and cancel the main call if they trip - the parallelization pattern applied to safety.
  • Enforce action rules in code at the tool boundary - ownership, amount limits, allowed recipients - not only in the prompt. Lab 13 found rules enforced by tools held in every prompt variant, while the one rule stated only in the prompt was broken in all of them.
  • Validate outputs that matter: schemas, citations present and valid, no internal data leaked.

Guardrail design and prompt-injection defences are covered in depth in Production Agents.

Gates and Human-in-the-Loop

GateTriggerExample
Action approvalSpecific tools, or arguments over a thresholdRefunds over $500; any external email
Plan approvalBefore executing a multi-step planA migration touching 40 files
Confidence gateA calibrated signal - verifier score, retrieval coverage, classifier probabilityLow grounding score โ†’ human review
EscalationPolicy ambiguity, user request, repeated failure"Speak to a person"

Two cautions. Model self-reported confidence ("I am 90% sure") is poorly calibrated; base gates on external signals - a verifier, test results, agreement between samples, or a trained classifier. And approval gates need the loop to pause and resume with persisted state; frameworks provide interrupts on top of checkpointing (Agents in Practice).

Clean handoff to a human means passing a summary, not a transcript: the customer's goal, what was verified, what was done, what failed and why, and the suggested next step.


Sub-Agents as Tools

An agent can call another model invocation as a tool: "research this question and return a 200-word summary with sources". The sub-agent runs its own loop in a fresh context, uses whatever tools it needs, and returns only a condensed result.

sequenceDiagram
    participant M as ๐Ÿค– Main agent
    participant S as ๐Ÿ”Ž Sub-agent (fresh context)
    participant T as ๐Ÿ”ง Tools

    M->>S: research(question, output_format)
    loop its own loop
        S->>T: search / read / run
        T-->>S: large raw results
    end
    S-->>M: condensed findings + sources (โ‰ˆ1K tokens)
    Note over M: main context grows by the summary,<br/>not by 40K tokens of raw pages
  • Why: context isolation. The main agent's context stays focused; exploration noise stays in the sub-agent. Coding agents use this for codebase searches; research agents for each line of inquiry.
  • Cost: the sub-agent re-reads whatever you pass it, and anything it doesn't return is lost to the main agent. Specify the task, the output format and the boundaries precisely - vague delegation is the most common failure (Anthropic, 2025).
  • Difference from multi-agent collaboration: a sub-agent-as-tool is a function call with a return value; the main agent stays in control and the sub-agent doesn't talk to peers. That keeps most of the single-agent simplicity.

Verification Loops That Work

The evaluator-optimizer pattern (Workflow Patterns) inside an agent: after producing a result, the agent (or harness) checks it and revises.

Feedback sourceStrengthExamples
ExecutionStrongest - objective and specificUnit tests, type checkers, linters, running a query, schema validation
Environment stateStrongRe-reading the record after an update; checking a file exists
Independent evaluator with new informationModerate - depends on the rubricCitation checker against sources; policy checklist
The same model re-reading its own outputWeak for reasoning errors (Huang et al., 2024)"Review your answer for mistakes"

Design for the strong end: give agents tools that produce feedback (a run_tests tool, a validate tool), and make "done" mean "verified". The module lab compares a no-feedback self-review loop with a test-feedback loop and best-of-n selection on the same problems.


Check Yourself

Check yourself
0 / 4 answered
  1. Why is model self-reported confidence a poor basis for a confidence gate?
  2. What does using a sub-agent as a tool mainly buy you?
  3. Which verification loop is most likely to fix a bug in generated code?
  4. What should a handoff to a human agent contain?

Exercises

Exercise - Guard an agent

For an agent that manages a user's calendar and email, list the guardrails you would add at input, action and output, and which actions need human approval. Justify each by the harm it prevents.

Solution

Input: scope check (calendar/email only), injection heuristics on email bodies. Action: sending email to external domains needs approval; bulk deletes blocked; invites to more than N people need approval; all writes scoped to the user's account. Output: no content from other users' threads; summaries cite message ids. Approval covers the irreversible, externally visible actions - sending and deleting - where a manipulated agent would do the most damage.

Exercise - Delegate precisely

Rewrite this delegation for a research sub-agent so its output is useful to the main agent: "Look into competitor pricing."

Solution

"Find the current list prices (monthly, per seat) for the team plans of Competitors A, B and C from their official pricing pages. Return a JSON array of {competitor, plan, price_usd, billing_period, source_url, retrieved_date}. If a price isn't public, set price_usd to null and say so. Don't include blog posts or third-party estimates. Stop after the three official pages."

Study Notes

  • Default: one agent, good tools, clear prompt; add structure only for observed limits
  • Guardrails: input, action (at the tool boundary, in code), output; run cheap checks in parallel
  • Gates: action approval, plan approval, confidence (external signals, not self-report), escalation; need pause/resume
  • Handoffs to humans: structured summaries
  • Sub-agents as tools: context isolation, precise delegation, higher total tokens; main agent stays in control
  • Verification: execution > environment state > independent evaluator with new info > self-review

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท