Contents
Map

16 · Agent Frameworks

Choosing an Agent Framework

View as:

Choosing an Agent Framework

An agent framework packages the pieces you built by hand in Code Lab 13-01 - the loop, tool schemas, state, approvals, tracing and multi-agent wiring - behind an API. This chapter maps what frameworks provide, how the major ones differ in their core abstraction, and how to choose one (or none) for a project.

Learning objectives 45 min
By the end of this page you will be able to:
  • List what an agent framework adds to a raw model API loop, and what it hides from you
  • Place a framework on the layer map - model SDK, agent SDK, orchestration runtime, managed platform
  • Compare LangGraph, OpenAI Agents SDK, Claude Agent SDK, Google ADK, CrewAI, Microsoft Agent Framework and PydanticAI by abstraction, state, human-in-the-loop, durability and model neutrality
  • Choose a framework for a given project with a written rationale, including when to use none

What a Framework Adds

Strip any framework down and you find the same loop: call the model with tools, execute the tool calls it returns, append the results, repeat until it answers or a bound trips. What frameworks add is everything around that loop:

ConcernHand-written loopWhat frameworks provide
Tool schemasWrite JSON schemas by handGenerated from type hints and docstrings; validation of arguments
The loop~50 lines you ownA runner with turn limits, error handling and retries
StateA list of messagesSessions, threads, checkpoints you can resume from
Human approvalYour own pause logicInterrupt-and-resume primitives (interrupt(), needs_approval, approval_mode, deferred tools)
Multi-agentYour own routingHandoffs, agents-as-tools, sequential/parallel/loop agents, graphs, group chat
Tracingprint statementsSpans per model call and tool call, often OpenTelemetry-compatible
IntegrationsYour own clientsModel providers, MCP clients, vector stores, deployment targets

What they hide is exactly what you debug: the prompt actually sent, how tool errors are turned into messages, how history is truncated, and when retries happen. Build the loop by hand once (Module 13) before you adopt a framework, so that you can read its traces.

The Layer Map

flowchart TD
    P["☁️ Managed agent platforms<br/>Agent Platform (Agent Runtime), Bedrock AgentCore,<br/>Foundry Agent Service, Claude Managed Agents"]
    O["🕸️ Orchestration runtimes<br/>LangGraph, MAF Workflows, ADK workflow agents,<br/>CrewAI Flows"]
    A["🤖 Agent SDKs<br/>OpenAI Agents SDK, Claude Agent SDK, Google ADK,<br/>PydanticAI, LangChain create_agent, MAF Agent"]
    M["🧠 Model SDKs<br/>openai, anthropic, google-genai"]
    P --> O --> A --> M

    style P fill:#ddd8e4,stroke:#b8b0c8
    style O fill:#e8e0d4,stroke:#c8b89a
    style A fill:#d8dfe8,stroke:#b0bac8
    style M fill:#dde4dc,stroke:#b0c4b0
  • Model SDKs give you a single call with tool definitions; you write the loop.
  • Agent SDKs give you the loop, tools, sessions and approvals for one agent, plus light multi-agent composition (handoffs, agents as tools).
  • Orchestration runtimes give you explicit graphs or workflows with checkpointed state - for long-running, branching or human-paced processes.
  • Managed platforms host and scale the agent, and supply sessions, memory, identity and observability as services.

Most frameworks span two layers: LangChain's create_agent is an agent SDK built on the LangGraph runtime; Google ADK and Microsoft Agent Framework have both agents and workflows.

The Frameworks at a Glance

Versions are those verified for this module's lab (September 2026).

Core abstractionState and durabilityHuman-in-the-loopMulti-agentModel neutrality
LangChain + LangGraph (1.4 / 1.2)Agent with middleware, on a state graphCheckpointers (memory, SQLite, Postgres); durability modes; long-term Storeinterrupt() + Command(resume=...); HumanInTheLoopMiddlewareGraphs, subgraphs, Send fan-out, supervisor/swarm librariesAny provider via integrations
OpenAI Agents SDK (0.22)Agent + RunnerSessions (SQLite, Redis, OpenAI Conversations); serialisable RunStateneeds_approval on tools; interruptions; RunState.approve/rejectHandoffs, agents as toolsOpenAI-first; other providers via Chat Completions or LiteLLM
Claude Agent SDK (0.2)The Claude Code harness as a librarySessions resumed by id, file checkpointingPermission modes, can_use_tool, hooksSubagentsClaude only
Google ADK (2.10)LlmAgent + workflow agents + RunnerSession services (memory, database, Agent Platform)Tool confirmation; callbacksSub-agents, Sequential/Parallel/LoopAgent, AgentTool, A2AGemini-first; others via LiteLLM
CrewAI (1.15)Role-playing agents in a crew; event-driven FlowsFlow state persistence; crew memoryhuman_input on tasks; Flow stepsSequential or hierarchical crews; FlowsAny provider via LiteLLM
Microsoft Agent Framework (1.19)Agent + graph-based WorkflowsSessions; workflow checkpointsapproval_mode on tools; request/response in workflowsSequential, concurrent, handoff, group chat, Magentic orchestrationsAzure/OpenAI-first; many connectors
PydanticAI (2.51)Typed agent with dependency injectionMessage history; durable execution via Temporal, DBOS or PrefectDeferred tools (requires_approval)Agents as tools, delegation, pydantic-graphAny provider

How to Choose

flowchart TD
    Q1{"Is the path<br/>predictable?"} -->|"yes"| W["Workflow in plain code<br/>(or a graph runtime if it is<br/>long-running)"]
    Q1 -->|"no"| Q2{"Need a coding / file-system<br/>agent with built-in tools?"}
    Q2 -->|"yes"| CL["Claude Agent SDK<br/>(or Codex-style harness)"]
    Q2 -->|"no"| Q3{"Long-running, branching,<br/>human-paced, must resume?"}
    Q3 -->|"yes"| G["Graph runtime:<br/>LangGraph / MAF Workflows<br/>+ durable execution"]
    Q3 -->|"no"| Q4{"Committed to one<br/>cloud or model vendor?"}
    Q4 -->|"Google"| ADK["Google ADK"]
    Q4 -->|"Microsoft"| MAF["Microsoft Agent Framework"]
    Q4 -->|"OpenAI"| OA["OpenAI Agents SDK"]
    Q4 -->|"no"| N["PydanticAI or LangChain<br/>create_agent"]

    style W fill:#dde4dc,stroke:#b0c4b0
    style CL fill:#d8dfe8,stroke:#b0bac8
    style G fill:#e8e0d4,stroke:#c8b89a
    style N fill:#ddd8e4,stroke:#b8b0c8

Questions that decide it in practice:

  1. Does the team already run on one cloud? The vendor framework gets you deployment, identity, observability and support in one place - ADK on Google Cloud's Agent Platform, Microsoft Agent Framework on Foundry, the OpenAI Agents SDK with OpenAI's hosted tools.
  2. How long does one run live? Seconds: any agent SDK. Hours to weeks with human steps: a checkpointed graph plus durable execution (Production Agents).
  3. How much control over the prompt and loop do you need? Frameworks that generate large hidden prompts (role/backstory templates, planning prompts) are harder to tune than thin ones; read the prompt a framework sends before committing.
  4. Can you swap the model? Keep tools as plain functions and business rules in the tools (as the lab does), and a framework change becomes a day's work.
  5. Is it maintained, and how fast is it changing? Most of these are pre-1.0 or recently past 1.0 and ship weekly. Pin versions and keep a small end-to-end test suite.

When to use no framework

A single model call with tools and a 50-line loop is often the better choice: fewer dependencies, a prompt you fully control, and nothing to upgrade. Anthropic's guidance is to start with the model API directly and add abstraction only when it demonstrably helps (Building Effective Agents, 2024). Use a framework when you need its durable state, approvals, tracing or integrations - not for the loop itself.

What the Lab Shows

Code Lab 16-01 builds the same shop agent - same tools, same policy prompt, same model, same grader - in seven frameworks. The plumbing differs (how tools are declared, how a refund pauses for approval and resumes, sync vs async runners). Most tasks behaved identically everywhere; the exception was a rule enforced only by the prompt, which three frameworks' agents followed every time and two never did. Replaying the requests traced the split to incidental tool-schema details the frameworks generate (JSON-schema title fields combined with additionalProperties: false), not to anything the framework does on purpose. Frameworks change the exact tokens the model sees, so test your tasks in the framework you pick - and enforce rules in code, where serialisation details can't move them.

Check Yourself

Check yourself
0 / 4 answered
  1. Which concern does an agent framework NOT remove from your responsibility?
  2. A process runs for three days with two human review steps and must survive restarts. Which layer do you need?
  3. Why build an agent loop by hand before adopting a framework?
  4. In the lab, two frameworks' agents cancelled another customer's order that three others refused. What caused it?

Exercises

Exercise - Choose and justify

For each project, choose a framework (or none) and write three sentences of rationale: (a) a support bot on Google Cloud that must hand off to a billing agent owned by another team; (b) a nightly job that summarises 200 documents into a report; (c) a loan-approval process with an underwriter review step that can take two days; (d) an internal coding assistant that edits repositories.

Solution

(a) Google ADK on Agent Platform; the billing agent exposed over A2A so the other team can use any framework. (b) No agent framework: a map-reduce workflow in plain code with batch API calls. (c) A graph runtime with checkpoints (LangGraph or MAF Workflows) on durable execution; the review is an interrupt that can wait days. (d) Claude Agent SDK (or another coding harness): built-in file and shell tools, permissions and hooks already exist.

Exercise - Read the hidden prompt

Pick two frameworks from the lab. Log the exact request body each sends to the model server for the same task (a logging proxy, or the framework's debug logging). Compare the system prompts and tool schemas. What did each framework add?

Hint

mlx_lm.server and vLLM can log requests; or put a small HTTP proxy between the script and the server

Solution

Thin SDKs (OpenAI Agents, PydanticAI, LangChain create_agent) send your instructions nearly verbatim plus the tool schemas. CrewAI wraps the backstory, role and goal in its own template with task and expected-output text; ADK adds agent name and transfer instructions when sub-agents exist. Schema details differ too (strict mode, additionalProperties). These differences are why the same model can behave differently across frameworks.

Study Notes

  • A framework = loop + schemas + state + approvals + multi-agent + tracing + integrations; it hides the prompt, error handling, truncation and retries
  • Layers: model SDK -> agent SDK -> orchestration runtime -> managed platform
  • Choose on cloud alignment, run length and durability, prompt control, model neutrality, and maintenance - not on accuracy claims
  • Keep tools as plain functions with business rules inside; pin versions; keep an end-to-end test suite
  • No framework is a legitimate choice for a simple loop

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·