Contents
Map

15 ยท Agent Patterns & Multi-Agent

Workflow Patterns

View as:

Workflow Patterns

Workflow patterns are the reusable shapes for composing model calls and tools when your code controls the sequence: chaining, routing, parallelization, orchestrator-workers and evaluator-optimizer. Most production "agentic" systems are built mainly from these, with a true agent loop only where the path can't be predicted. This chapter gives each pattern's structure, when it pays off, its costs and its failure modes.

Learning objectives 50 min
By the end of this page you will be able to:
  • Describe the five workflow patterns - prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer - and draw each
  • Choose a pattern (or combination) for a task from its structure, latency budget and verification options
  • Add gates and checks between steps so errors don't propagate
  • Estimate the cost and latency of a composed workflow before building it
Prerequisites

The Building Block: an Augmented Model Call

Every pattern composes one unit - a model call that may use retrieval, tools and memory. Anthropic's taxonomy (Building Effective Agents, 2024) names five compositions of that unit; each trades extra calls for accuracy, speed or control.

flowchart TD
    subgraph C["๐Ÿ”— Prompt chaining"]
        C1["Step 1"] --> CG{"Gate"} --> C2["Step 2"] --> C3["Step 3"]
    end
    subgraph R["๐Ÿ”€ Routing"]
        R0["Classify"] --> R1["Handler A"]
        R0 --> R2["Handler B"]
        R0 --> R3["Handler C"]
    end
    subgraph P["โšก Parallelization"]
        P0["Split"] --> P1["Call 1"] & P2["Call 2"] & P3["Call 3"] --> PA["Aggregate"]
    end
    subgraph O["๐Ÿงญ Orchestrator-workers"]
        O0["Orchestrator plans"] --> O1["Worker"] & O2["Worker"] --> OS["Orchestrator synthesises"]
    end
    subgraph E["๐Ÿ” Evaluator-optimizer"]
        E1["Generate"] --> E2{"Evaluate"}
        E2 -->|"revise with feedback"| E1
        E2 -->|"accept"| E3["Output"]
    end

    style C fill:#e8e2d9,stroke:#ccc4b8
    style R fill:#d8dfe8,stroke:#b0bac8
    style P fill:#dde4dc,stroke:#b0c4b0
    style O fill:#e8e0d4,stroke:#c8b89a
    style E fill:#ddd8e4,stroke:#b8b0c8

1. Prompt Chaining

A fixed sequence of calls, each consuming the previous output: outline โ†’ draft โ†’ edit; extract โ†’ validate โ†’ transform. Between steps you can insert programmatic gates - schema validation, a length check, a test - that stop or repair bad intermediate output before it propagates.

  • Use when the task decomposes cleanly into fixed steps, and each step is easier for the model than the whole.
  • Gains accuracy per step, debuggability (inspect every intermediate), and the ability to use a cheaper model for easy steps.
  • Costs latency adds up linearly; errors compound if gates are weak.

2. Routing

A classifier (a model call, a small fine-tuned model, or rules) sends each input to a specialised handler: refunds to the refund flow, technical questions to the docs-RAG flow, easy questions to a small model and hard ones to a large model.

  • Use when inputs fall into distinct categories that are better handled separately, or when cost varies a lot by difficulty.
  • Gains specialised prompts that don't interfere, and cost savings from sending easy traffic to cheap models.
  • Costs misroutes are silent errors - evaluate the router as its own component (a confusion matrix on labelled inputs) and give it a fallback route.

3. Parallelization

Run calls concurrently and aggregate, in two flavours:

FlavourHowExample
SectioningSplit a task into independent parts handled in parallelReview a contract clause by clause; run a guardrail check alongside the main response
VotingRun the same task several times and aggregateMajority vote on a classification (self-consistency, Wang et al. 2023); best-of-n with a verifier

Latency is the slowest branch rather than the sum. Voting improves reliability when errors are independent and there is a good aggregation rule; it is weak when all samples share the same misconception. Best-of-n is only as good as its selector: with an executable verifier (tests, a schema, a checker) it is powerful; with the same model picking its favourite it helps much less.

4. Orchestrator-Workers

An orchestrator model decides at run time how to split the task, dispatches sub-tasks to workers, and synthesises their results. Unlike sectioning, the sub-tasks aren't known in advance - which files to change, which sources to research.

  • Use when the decomposition depends on the input (multi-file code changes, open research questions).
  • Gains parallel exploration and focused worker contexts.
  • Costs more tokens (each worker re-reads context), and the orchestrator's plan quality caps the result. This is the step from workflows toward multi-agent systems (Multi-Agent Architectures).

5. Evaluator-Optimizer

One call generates, another evaluates against criteria and returns feedback, and the loop repeats until the evaluator accepts or a round limit is hit.

  • Use when there are clear evaluation criteria and iteration measurably helps: code against tests, translation against a glossary, output against a schema or rubric.
  • The evaluator must add information. Feedback from execution (tests, validators, a compiler) or from a rubric the generator didn't see works. A second call of the same model re-reading the same material with nothing new - "review your answer" - is weak on reasoning tasks (Huang et al., 2024), and on modern reasoning models it mostly repeats thinking already done. The module's lab measures the difference on coding problems.
  • Always bound the loop (2-3 rounds is typical) and keep the best-scoring output, not just the last.

Choosing and Combining

Task shapePattern
Known, ordered stepsChaining, with gates
Distinct input categories, or wide difficulty rangeRouting
Independent parts, or reliability by redundancyParallelization (sectioning / voting)
Unknown decompositionOrchestrator-workers
Checkable quality criteriaEvaluator-optimizer
Unknown path and many dependent stepsAn agent loop (Agent Foundations)

Real systems combine them: a support bot routes by intent; the refund route chains extraction โ†’ policy check โ†’ action; a guardrail runs in parallel; the research route is an orchestrator-workers flow whose report goes through an evaluator-optimizer citation check.

Estimating cost and latency

Write the call graph down before building. For each node: expected input and output tokens, model, and whether it is on the critical path. Cost is the sum over all nodes (times the expected number of loop iterations); latency is the longest path. Example: route (small model, 0.3 s) โ†’ three parallel section reviews (large model, 6 s each) โ†’ synthesis (4 s) โ†’ evaluator-optimizer with up to 2 rounds (2 ร— 5 s): latency โ‰ˆ 0.3 + 6 + 4 + 10 โ‰ˆ 20 s worst case; cost โ‰ˆ 1 + 3 + 1 + 2 ร— 2 = 9 calls' worth of tokens.


Check Yourself

Check yourself
0 / 4 answered
  1. What distinguishes orchestrator-workers from parallelization by sectioning?
  2. When is an evaluator-optimizer loop most likely to improve results?
  3. A router sends 5% of refund requests to the general-FAQ handler. How would you find and fix this?
  4. Best-of-5 with the same model choosing its favourite gives little gain; best-of-5 selected by unit tests gives a large one. Why?

Exercises

Exercise - Design a pipeline

Design the workflow for turning 200-page supplier contracts into a risk summary: identify the patterns you'd use, draw the call graph, add gates, and estimate calls and latency per contract.

Solution

Chunk by clause headings; parallel sectioning: one call per clause group extracts obligations, liabilities and termination terms into a schema (gate: schema validation, retry once); routing: clauses flagged high-risk go to a larger model for detailed analysis; synthesis call writes the summary with clause citations; evaluator-optimizer: a checker verifies every claim cites a clause id (programmatic) and a rubric call checks coverage of the risk checklist (max 2 rounds). With ~40 clause groups at 3 s in parallel batches of 10: ~12 s; high-risk analysis ~8 s; synthesis ~10 s; evaluation up to 2 ร— 6 s - roughly 40-45 s and ~50 calls per contract.

Exercise - Voting arithmetic

A classifier is right 80% of the time and its errors are independent across samples. What is the accuracy of a majority vote of 3? Of 5? Why is the real gain usually smaller?

Solution

3 votes: P(at least 2 right) = 0.8ยณ + 3 ร— 0.8ยฒ ร— 0.2 = 0.512 + 0.384 = 0.896. 5 votes: ฮฃ_{kโ‰ฅ3} C(5,k) 0.8^k 0.2^(5-k) = 0.328 + 0.410 + 0.205 = 0.942. Real samples from one model share systematic errors (the same misreading of an ambiguous input), so errors are correlated and the gain is smaller - sometimes near zero on the hard cases that matter.

Study Notes

  • Unit: an augmented model call (retrieval, tools, memory)
  • Chaining: fixed steps + programmatic gates between them
  • Routing: classify then specialise; evaluate the router; fallback route
  • Parallelization: sectioning (independent parts) and voting (redundancy); best-of-n needs a real verifier
  • Orchestrator-workers: run-time decomposition; bridge to multi-agent
  • Evaluator-optimizer: works when the evaluator adds information (tests, rubric); bound rounds; keep the best output
  • Combine patterns; estimate cost (sum of calls) and latency (longest path) before building

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท