Contents
Map

17 ยท Production Agents

Production Agent Architecture

View as:

Production Agent Architecture

A production agent is a distributed system with a model in the middle: requests arrive through an interface, a durable orchestrator runs the loop, tools reach real systems under least privilege, humans approve what matters, and everything is traced and budgeted. This chapter lays out that architecture and the reliability engineering that makes it survive real traffic; the following chapters go deep on security, durability, cost, evaluation and observability.

Learning objectives 50 min
By the end of this page you will be able to:
  • Draw the layers of a production agent system and assign each responsibility (auth, orchestration, tools, state, approvals, telemetry) to a layer
  • Set bounds, timeouts, retries, idempotency keys and circuit breakers for an agent's model and tool calls
  • Design human-in-the-loop gates - what triggers them, what the reviewer sees, what happens on timeout - and their audit trail
  • Choose synchronous, streaming or asynchronous execution for an agent workload and scale it with queues

The Layers

flowchart TD
    I["๐Ÿšช Interface<br/>API / chat UI / webhooks, authentication, input limits"]
    O["๐Ÿงญ Orchestrator (durable)<br/>agent loop or workflow, state, bounds, HITL gates"]
    M["๐Ÿง  Model gateway<br/>routing, fallbacks, caching, rate limits, budgets"]
    T["๐Ÿ”ง Tool layer<br/>APIs, MCP servers, code sandbox - scoped credentials, policy checks"]
    S["๐Ÿ’พ State and memory<br/>run state, conversation, long-term memory, artifacts"]
    H["๐Ÿ‘ฉโ€โš–๏ธ Human review<br/>approval queue, audit log"]
    X["๐Ÿ“Š Telemetry<br/>traces, metrics, evals, cost"]
    I --> O
    O <--> M
    O <--> T
    O <--> S
    O <--> H
    O -.-> X
    M -.-> X
    T -.-> X

    style O fill:#d8dfe8,stroke:#b0bac8
    style M fill:#ddd8e4,stroke:#b8b0c8
    style T fill:#e8e0d4,stroke:#c8b89a
    style H fill:#dde4dc,stroke:#b0c4b0
    style X fill:#e8e2d9,stroke:#ccc4b8
LayerResponsibilitiesChapter
InterfaceAuthenticate the user and carry their identity into the run; size limits; accept a task and return an id for long runs-
OrchestratorRun the loop or workflow; persist state after every step so a crash doesn't lose progress; enforce bounds; pause for approvalsDurable Execution
Model gatewayOne place for provider keys, routing between models, fallbacks, prompt caching, rate limits and per-tenant budgetsCost & Latency
Tool layerExecute tool calls with the user's scoped credentials; policy checks in code; sandbox for code and browsingAgent Security
State and memoryRun checkpoints, conversation history, long-term memory, artifacts (files the agent produced)Agent Memory
Human reviewApproval queue with full context; decisions recordedBelow
TelemetryA trace per run, spans per model and tool call, cost and outcome metrics, online evalsObservability, Evaluation

Two principles cut across the layers. Enforce policy outside the model: the model proposes actions, code decides whether they're allowed (Lab 13 and Lab 14 measured why). Make every step recoverable: persisted state, idempotent writes and replayable steps turn a crash or a timeout into a retry instead of an incident.

Reliability Engineering

Bounds on everything

An agent loop without limits runs until it exhausts a quota. Bound each run:

@dataclass
class RunLimits:
    max_model_calls: int = 25
    max_tool_calls: int = 40
    max_wall_seconds: int = 300
    max_cost_usd: float = 1.00
    max_repeated_call: int = 3          # same tool + same arguments
    max_context_tokens: int = 150_000   # compact or summarise above this

When a bound trips, stop cleanly: persist what was done, return a partial result with an explicit status ("limit_reached"), and emit a metric - a rising rate of limit hits is an early signal that a model, prompt or tool changed.

Timeouts, retries and idempotency

Classify failures before retrying:

FailureAction
Transient (timeout, 429, 5xx, overloaded)Retry with exponential backoff and jitter; honour retry-after
Invalid request or argumentsDon't retry blindly - return the error to the model as a tool result so it can correct
Permanent (not found, forbidden)Don't retry; return the error or escalate
Model refusal or content filterDon't retry the same request; route to a human or a safe fallback

Retries are only safe if writes are idempotent: send an idempotency key (run id + step id) with every external write, and make the tool check whether that key was already applied. Durable-execution engines give you the step ids for free.

Circuit breakers stop hammering a failing dependency: after N consecutive failures, fail fast for a cool-down period, then let a probe request through. For model providers, the fallback is usually another model or region behind the gateway.

Graceful degradation

Decide in advance what the agent does when a dependency is down: answer from cache, switch to read-only mode, queue the write for later, or hand off to a human with the context gathered so far. "The agent errors" is the worst option; "the agent claims success" is worse still.

Human-in-the-Loop Design

Human review is a design element with its own requirements, not a fallback.

When to require approval - rules in code, not the model's judgement:

def needs_approval(call: ToolCall, ctx: RunContext) -> bool:
    return (
        call.name in IRREVERSIBLE                   # payments, deletions, external messages
        or call.amount_usd > ctx.tenant.approval_threshold
        or call.affects_records > 10                # bulk operations
        or call.recipient_domain not in ctx.tenant.trusted_domains
    )

What the reviewer sees: the exact action and arguments (the email as it will be sent, not a summary), why the agent wants it (the relevant part of the trace), what happens on approve and on reject, and an edit option. Reviewers approve what they are shown; showing a summary invites rubber-stamping.

Timeouts: a pending approval needs a policy - remind, escalate, then cancel the run (not approve by default). The orchestrator must be able to wait hours or days without holding a process, which is why approvals and durable execution go together.

Audit: record who decided, what they saw, when, and what happened next. Regulated domains require it; debugging does too.

Rubber-stamping is the failure mode: if 99% of requests are approved unread, the gate provides no safety. Keep the approval rate meaningful by narrowing what needs approval, and sample approved actions for review.

Execution Modes and Scaling

ModeUse whenMechanism
SynchronousAnswers in a few secondsRequest/response
StreamingInteractive chat; long answersSSE or WebSocket; stream tokens and progress events (tool started, tool finished)
AsynchronousMinutes to days; approvals; batchAccept, return a task id, run on workers from a queue, notify by webhook or poll (A2A tasks and MCP Tasks use the same shape)

Agent workers are mostly waiting on model and tool I/O, so scale them on queue depth and concurrent runs, not CPU. Put rate limits at the gateway per provider and per tenant; the provider's tokens-per-minute limit is usually the real ceiling. Keep workers stateless - all run state in the durable store - so any worker can pick up any step.

Check Yourself

Check yourself
0 / 4 answered
  1. An agent's refund tool is retried after a network timeout and the customer is refunded twice. What was missing?
  2. A tool call fails with a validation error on its arguments. What should the harness do?
  3. What should happen when an approval request gets no answer for 24 hours?
  4. Why scale agent workers on queue depth rather than CPU?

Exercises

Exercise - Design the layers

An agent handles expense reports: it reads receipts, checks them against policy, files the report in the finance system and emails the employee. Assign every responsibility to a layer, list the bounds you'd set, the actions needing approval, and what happens when the finance API is down.

Solution

Interface: SSO identity, file-size limits, async task id. Orchestrator: durable workflow (extract -> check -> file -> notify) with bounds (e.g. 15 model calls, $0.50, 10 min). Tools: receipt OCR, policy lookup, finance API with the employee's scoped token, email limited to the employee's address. Approvals: reports over the threshold or with policy exceptions go to the manager with the receipts and the flagged rule. Finance API down: circuit breaker opens, the workflow waits and retries the idempotent file step later, the employee is told it is queued - never "filed".

Exercise - Failure classes

Take ten failed runs from Lab 13 or your own agent and classify each failure as transient, invalid arguments, permanent, refusal, or model behaviour (wrong tool, claimed action). Which are fixed by retries, and which need a different fix?

Solution

Retries fix only the transient class. Invalid arguments need the error returned to the model (and better schemas); permanent errors need escalation; model-behaviour failures need tool or harness changes (guards, state checks) and evaluation - retrying a claimed action produces the same claim.

Study Notes

  • Layers: interface, durable orchestrator, model gateway, tool layer, state and memory, human review, telemetry
  • Policy outside the model; every step recoverable
  • Bound model calls, tool calls, time, cost, repeats, context; stop with an explicit partial status
  • Classify failures; retry only transient ones; idempotency keys on writes; circuit breakers; planned degradation
  • HITL: rules in code, exact actions shown, timeout policy (never auto-approve), audit, watch for rubber-stamping
  • Sync / streaming / async; scale on queue depth; rate limits at the gateway

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท