Contents
Map

16 ยท Agent Frameworks

OpenAI Agents SDK

View as:

OpenAI Agents SDK

The OpenAI Agents SDK is a small, Python-first agent library with a handful of primitives - agents, tools, handoffs, guardrails, sessions - and a Runner that runs the loop, with tracing built in. It is the production successor to OpenAI's experimental Swarm, released in March 2025 alongside the Responses API.

Learning objectives 40 min
By the end of this page you will be able to:
  • Build an agent with function tools and run it with Runner, synchronously or asynchronously
  • Compose agents with handoffs and agents-as-tools, and explain the difference
  • Add input and output guardrails, tool approvals with resumable RunState, and sessions
  • Run the SDK against a non-OpenAI model and know what you lose

Primitives

flowchart LR
    U["๐Ÿ‘ค Input"] --> IG["๐Ÿ›ก๏ธ Input guardrails"]
    IG --> A["๐Ÿค– Agent<br/>instructions, tools, model"]
    A -->|"tool call"| T["๐Ÿ› ๏ธ Function / hosted / MCP tools"]
    T --> A
    A -->|"handoff"| B["๐Ÿค– Specialist agent"]
    A --> OG["๐Ÿ›ก๏ธ Output guardrails"] --> O["๐Ÿ’ฌ final_output"]
    R["๐Ÿ” Runner<br/>loop, max_turns, tracing"] -.-> A
    S["๐Ÿ’พ Session<br/>conversation history"] -.-> R

    style A fill:#d8dfe8,stroke:#b0bac8
    style B fill:#d8dfe8,stroke:#b0bac8
    style IG fill:#e8e0d4,stroke:#c8b89a
    style OG fill:#e8e0d4,stroke:#c8b89a
    style R fill:#dde4dc,stroke:#b0c4b0
PrimitiveWhat it is
AgentInstructions, tools, handoffs, guardrails, optional output_type (a Pydantic model for structured output), model and ModelSettings
function_toolA decorator turning a typed Python function into a tool; the schema comes from type hints and the docstring, strict by default
Hosted toolsWebSearchTool, FileSearchTool, code interpreter, HostedMCPTool - run by OpenAI with the Responses API
MCPMCPServerStdio, MCPServerStreamableHttp to use any MCP server's tools
HandoffTransfer the conversation to another agent, which then owns the reply
GuardrailsInput guardrails run on the user input (in parallel with the agent by default); output guardrails check the final output; a tripwire raises an exception
Runnerrun (async), run_sync, run_streamed; max_turns bounds the loop
SessionStores history across runs: SQLiteSession, Redis, OpenAIConversationsSession

An Agent with Handoffs, a Guardrail and an Approval

from pydantic import BaseModel
from agents import (Agent, GuardrailFunctionOutput, Runner, SQLiteSession, function_tool,
                    input_guardrail)

@function_tool(needs_approval=True)
def refund_item(order_id: str, item_id: int) -> dict:
    """Refund one line item of a delivered order."""
    ...

class Topic(BaseModel):
    off_topic: bool

classifier = Agent(name="classifier", instructions="Is this message unrelated to orders?",
                   output_type=Topic)

@input_guardrail
async def on_topic(ctx, agent, user_input):
    result = await Runner.run(classifier, user_input, context=ctx.context)
    return GuardrailFunctionOutput(output_info=result.final_output,
                                   tripwire_triggered=result.final_output.off_topic)

billing = Agent(name="billing", instructions="Handle refunds.", tools=[refund_item])
triage = Agent(name="triage", instructions="Route the customer to the right specialist.",
               handoffs=[billing], input_guardrails=[on_topic])

session = SQLiteSession("customer-42", "conversations.db")
result = Runner.run_sync(triage, "Refund the trekking poles on O1006, please.", session=session)
while result.interruptions:                      # refund_item paused for approval
    state = result.to_state()                    # serialisable: state.to_json() survives a restart
    for item in result.interruptions:
        state.approve(item)                      # or state.reject(item)
    result = Runner.run_sync(triage, state, session=session)
print(result.final_output, result.last_agent.name)

Handoffs vs agents as tools. A handoff passes control: the specialist sees the conversation and answers the user directly (result.last_agent tells you who finished). agent.as_tool(...) keeps control with the caller: the specialist runs on a generated input and returns a result the caller uses. Use handoffs for triage into departments; agents-as-tools for an orchestrator that combines specialists' outputs (the manager pattern from Multi-Agent Architectures).

Approvals. needs_approval accepts True or an async function of the call's arguments, so you can require approval only above an amount. The run stops with result.interruptions; to_state() gives a RunState you can serialise, store and resume later - a human can approve tomorrow.

Tracing

Every run records a trace: spans for agent runs, model calls, tool calls, handoffs and guardrails. By default traces are uploaded to the OpenAI dashboard; add processors to export them elsewhere (many observability vendors integrate), or call set_tracing_disabled(True) - the lab does, because it runs against a local server. Traces are the first thing to open when an agent misbehaves.

Other Models

The SDK defaults to the OpenAI Responses API. To use any OpenAI-compatible server (vLLM, mlx_lm.server, Ollama) wrap an AsyncOpenAI client in OpenAIChatCompletionsModel, as the lab does; a LiteLLM extension covers other providers. Hosted tools (web search, file search, hosted MCP) and some tracing detail depend on the Responses API, so they don't work through Chat Completions models.

When to Use It

Good fit: OpenAI-centred stacks; triage-and-specialist designs; teams that want few concepts and built-in tracing. Less good: long-running processes with explicit branching and multi-day waits (combine with a durable-execution engine, or use a graph runtime), or teams that must stay model-neutral and want hosted tools.

Check Yourself

Check yourself
0 / 4 answered
  1. The triage agent hands off to billing. Who writes the reply the user sees?
  2. How does a run requiring approval survive a process restart?
  3. Why does the lab call set_tracing_disabled(True)?
  4. An input guardrail classifies every message with a second model call. What does it cost, and how does the SDK hide the latency?

Exercises

Exercise - Approval above a threshold

In the lab's agent_openai.py, replace needs_approval=True with an async function that requires approval only when the refunded item is worth more than 50.00. Log which refunds were approved automatically.

Hint

The function receives (run_context, args_dict, call_id)

Hint

Look the item up in the shop to get its price

Solution

async def over_50(ctx, args, call_id): order = shop.get_order(args["order_id"]); item = next(i for i in order["items"] if i["item_id"] == args["item_id"]); return item["qty"] * item["unit_price"] > 50. Pass needs_approval=over_50 for refund_item. Small refunds run straight through; large ones appear in result.interruptions.

Exercise - Split into triage and specialists

Rebuild the shop agent as a triage agent with handoffs to an orders agent (status, cancel, address) and a refunds agent (refund_item). Run the lab's six tasks three times each and compare with the single agent. Did the split help?

Solution

Typically not on tasks this small: the handoff adds a turn and a chance to route wrongly, and each specialist sees fewer tools but the same difficulty. Report task success and tokens per task; the single agent usually matches or beats the split, consistent with the evidence in Module 15.

Study Notes

  • Primitives: Agent, function tools (+ hosted and MCP tools), handoffs, guardrails, sessions, Runner, tracing
  • Handoff = transfer control; as_tool = call and return
  • needs_approval -> result.interruptions -> to_state() -> approve/reject -> resume (serialisable)
  • Tracing on by default to the OpenAI dashboard; disable or re-route for other backends
  • Non-OpenAI models via OpenAIChatCompletionsModel or LiteLLM; hosted tools need the Responses API

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท