17 - Production Agents
A demo agent needs to work once; a production agent needs to work reliably, safely and affordably on inputs nobody anticipated, for users who can be attacked through it. This module covers the engineering around the agent - architecture and reliability, security, durable execution, cost and latency, evaluation and benchmarks, and observability - and applies it in four worked system designs.
Learning objectives 9-11 hours
By the end of this module you will be able to:- Design a production agent system end to end - layers, bounds, retries, idempotency, human approval and execution mode
- Threat-model an agent with the lethal trifecta and the OWASP Agentic Top 10, and choose structural defences
- Make long-running agents durable and safe to retry with a workflow engine or checkpoints
- Model and reduce an agent's cost and latency without losing quality
- Run evaluation as a practice (suites, validated judges, CI and online) and read public agent benchmarks critically
- Instrument agents with OpenTelemetry GenAI traces and debug from them
Prerequisites
Where This Module Fits
flowchart LR
A["01 Architecture &<br/>reliability"] --> S["02 Security"]
A --> D["03 Durable<br/>execution"]
A --> C["04 Cost &<br/>latency"]
A --> E["05 Evaluation &<br/>benchmarks"]
E --> O["06 Observability"]
S & D & C & O --> SD["🏗️ System designs"]
style A fill:#d8dfe8,stroke:#b0bac8
style S fill:#e8e0d4,stroke:#c8b89a
style D fill:#dde4dc,stroke:#b0c4b0
style C fill:#e8e2d9,stroke:#ccc4b8
style E fill:#ddd8e4,stroke:#b8b0c8
style SD fill:#dde4dc,stroke:#b0c4b0
Chapter Map
| # | Chapter | You will learn | Time |
|---|---|---|---|
| 1 | Production Agent Architecture | Layers, bounds, failure classes and retries, idempotency, circuit breakers, HITL design, execution modes and scaling | 50 min |
| 2 | Agent Security | Prompt injection, lethal trifecta, OWASP Agentic Top 10, injection-resistant patterns and CaMeL, sandboxing and credentials | 55 min |
| 3 | Durable Execution | Replay and determinism, Temporal/Restate/DBOS vs checkpoints, durable approvals, an agent loop as a workflow | 45 min |
| 4 | Cost and Latency | Token cost model, caching, context hygiene, routing and cascades, batch, latency levers, budgets | 45 min |
| 5 | Evaluation and Benchmarks | Eval suites, task sourcing, validated judges, CI and online evals; SWE-bench Verified, τ-bench, GAIA, WebArena, OSWorld, Terminal-Bench, BFCL, AgentDojo | 50 min |
| 6 | Agent Observability | OpenTelemetry GenAI traces, content capture policy, dashboards and alerts, debugging from traces | 40 min |
| 7 | Q&A Review Bank | 32 questions across the module | 45 min |
System Designs
Interview-style designs, each ending with a Production Controls section that maps it onto the chapters above.
| # | Case study | Domain |
|---|---|---|
| 1 | Insurance Claims Processor | Multimodal evidence, fraud checks, adjuster approval |
| 2 | Prior Authorization | Healthcare workflow with policy retrieval and SLAs |
| 3 | Smart Diagnostic Assistant 1 | Field-service assistant with offline and safety constraints |
| 4 | Smart Diagnostic Assistant 2 | Orchestrator-subagent alternative with a latency budget |
Mini-Project
Take the MCP shop agent from Lab 14 to production readiness:
- Run it as a durable workflow with a refund approval that can wait a day, and idempotent writes.
- Add bounds, budgets, and OpenTelemetry traces exported to a local backend.
- Build regression, capability and safety suites (including Lab 14's result injection) and a CI gate.
- Write a one-page threat model (trifecta, OWASP ASI mapping) and a cost model per resolved ticket.
Review
- Q&A Review Bank - consolidated questions for this module
- Module quiz - every Check Yourself question in this module, in course order
Previous: 16 - Agent Frameworks · Next: 18 - Agent Engineering
Section Appendix
Summary & Key Terms - a quick recap of this section and its essential vocabulary.