Contents
Map

13 ยท Agent Foundations

The Agent Loop

View as:

The Agent Loop

The agent loop is the few dozen lines of code that turn a model into an agent: call the model, run the tools it asks for, send back the results, repeat until it answers. This chapter writes the loop correctly for the two main API shapes, lists the invariants that break it when violated, and adds the bounds and failure handling every production loop needs.

Learning objectives 60 min
By the end of this page you will be able to:
  • Write a working tool-use loop against the Claude Messages API and the OpenAI Responses API
  • State the loop invariants - every tool call answered, history kept unmodified, results in one turn - and what breaks when each is violated
  • Handle every stop reason, including truncation, refusal and paused server-side turns
  • Add bounds, repeated-call detection and error results, and run independent tool calls in parallel
Prerequisites

The Loop in One Picture

flowchart TD
    S(["๐ŸŽฏ Goal + instructions + tools"]) --> M["๐Ÿง  Call the model<br/>with the full history"]
    M --> D{"Stop reason?"}
    D -->|"tool calls"| X["๐Ÿ”ง Validate and execute<br/>every tool call"]
    X --> A["๐Ÿ“Ž Append the assistant turn<br/>and all results"]
    A --> B{"Within bounds?"}
    B -->|"yes"| M
    B -->|"no"| P(["โน๏ธ Stop: partial result<br/>+ status"])
    D -->|"final answer"| F(["โœ… Return the answer"])
    D -->|"truncated / refused /<br/>context full"| E(["โš ๏ธ Handle explicitly"])

    style M fill:#e8e0d4,stroke:#c8b89a
    style X fill:#dde4dc,stroke:#b0c4b0
    style A fill:#d8dfe8,stroke:#b0bac8
    style F fill:#dde4dc,stroke:#b0c4b0
    style E fill:#ddd8e4,stroke:#b8b0c8

The model is stateless: between calls it remembers nothing. What looks like memory is the growing message list your loop sends on every call. That single fact explains most of agent engineering - context growth, cost growth, and why the order and integrity of the history matter.


The Loop Against the Claude Messages API

import json
import anthropic

client = anthropic.Anthropic()

TOOLS = [{
    "name": "get_weather",
    "description": "Get the current weather for a city. Call it whenever the user asks about weather.",
    "input_schema": {
        "type": "object",
        "properties": {"city": {"type": "string", "description": "City name, e.g. Paris"}},
        "required": ["city"],
    },
}]

def run_tool(name: str, args: dict) -> str:
    if name == "get_weather":
        return json.dumps({"city": args["city"], "temp_c": 21, "sky": "clear"})
    raise ValueError(f"unknown tool {name}")

def run_agent(user_msg: str, max_steps: int = 10) -> str:
    messages = [{"role": "user", "content": user_msg}]
    for _ in range(max_steps):
        resp = client.messages.create(
            model="claude-opus-5", max_tokens=16000, tools=TOOLS, messages=messages)

        if resp.stop_reason == "refusal":                      # check before reading content
            return "The request was declined."
        if resp.stop_reason == "max_tokens":                   # a tool_use block may be cut off
            raise RuntimeError("Output truncated - raise max_tokens; don't run partial tool calls")
        if resp.stop_reason == "model_context_window_exceeded":
            raise RuntimeError("Context full - compact the history and retry")

        # Keep the assistant turn exactly as returned (text, thinking and tool_use blocks).
        messages.append({"role": "assistant", "content": resp.content})

        if resp.stop_reason == "pause_turn":                   # a server-side tool loop paused
            continue                                           # resend as-is; the server resumes
        if resp.stop_reason != "tool_use":                     # end_turn / stop_sequence
            return "".join(b.text for b in resp.content if b.type == "text")

        results = []
        for block in resp.content:
            if block.type != "tool_use":
                continue
            try:
                out, is_error = run_tool(block.name, block.input), False
            except Exception as e:                             # report failures to the model
                out, is_error = f"Error: {e}", True
            results.append({"type": "tool_result", "tool_use_id": block.id,
                            "content": out, "is_error": is_error})
        messages.append({"role": "user", "content": results})  # all results in ONE user turn
    return "Stopped: step limit reached."

Anthropic's SDKs also provide a tool runner (client.beta.messages.tool_runner) that runs this loop for you from decorated Python functions, with hooks for approval and result modification. Write the loop by hand once anyway - it is what every framework wraps.

Stop reasons and what to do

stop_reasonMeaningLoop action
tool_useThe model wants one or more tools runExecute all, send all results back, call again
end_turnThe model has finishedReturn the text
max_tokensOutput hit the limit - a tool call may be incompleteDon't execute; retry with more tokens
pause_turnA server-side tool loop (web search, code execution) reached its iteration limitAppend the assistant turn unchanged and call again - no extra "continue" message
refusalDeclined by the model or a safety classifierStop; surface it; optionally retry on a fallback model
model_context_window_exceededInput plus output filled the context windowCompact or trim history, then retry
stop_sequenceA custom stop sequence was producedApplication-specific

The Loop Against the OpenAI Responses API

The Responses API represents the conversation as a list of items. A tool call is a function_call item with a call_id; you answer it with a function_call_output item carrying the same call_id.

import json, os
from openai import OpenAI

client = OpenAI()
MODEL = os.environ["OPENAI_MODEL"]           # pick a current model - see Model Landscape

TOOLS = [{
    "type": "function", "name": "get_weather",
    "description": "Get the current weather for a city.",
    "parameters": {"type": "object", "properties": {"city": {"type": "string"}},
                   "required": ["city"], "additionalProperties": False},
    "strict": True,
}]

def run_agent(user_msg: str, max_steps: int = 10) -> str:
    items = [{"role": "user", "content": user_msg}]
    for _ in range(max_steps):
        resp = client.responses.create(model=MODEL, input=items, tools=TOOLS)
        calls = [it for it in resp.output if it.type == "function_call"]
        if not calls:
            return resp.output_text
        items += resp.output                   # includes reasoning items - they must be passed back
        for call in calls:
            try:
                out = run_tool(call.name, json.loads(call.arguments))
            except Exception as e:
                out = json.dumps({"error": str(e)})
            items.append({"type": "function_call_output", "call_id": call.call_id, "output": out})
    return "Stopped: step limit reached."

Instead of resending the items, you can pass previous_response_id=resp.id and send only the new function_call_output items; the server keeps the history. Most open-source servers (vLLM, SGLang, Ollama, mlx-lm) expose the older Chat Completions shape, where tool calls arrive in message.tool_calls and results go back as {"role": "tool", "tool_call_id": ..., "content": ...} messages - that is the shape the lab uses.


Loop Invariants

These rules are where hand-written loops break. Violating the first three produces API errors; violating the rest produces subtle quality loss.

InvariantWhy
Every tool call gets exactly one result, matched by id (tool_use_id / call_id / tool_call_id)An unanswered call makes the next request invalid; a duplicated one confuses the model
Send all results for one turn together - in Claude's API, in a single user message with the tool_result blocks firstSplit results break the call/result pairing the model was trained on
Append the assistant turn unmodified - including thinking blocks (Claude) and reasoning items (OpenAI)Reasoning models continue their reasoning across tool calls; stripped or edited reasoning is rejected or degrades the next step
Don't edit or reorder earlier turns once sentBreaks prompt caching, and on recent models invalidates preserved reasoning
Return errors as results, marked as errors (is_error: true on Claude)The model can recover from an informative error; it can't recover from a crashed loop
Never execute a truncated or refused turn's tool callsA cut-off argument often still parses as valid JSON

What the Model Sees at Each Step

For "What's the weather in Tokyo and Paris?", a capable model emits two tool calls in one turn (parallel tool use). The history after one iteration:

user:       What's the weather in Tokyo and Paris?
assistant:  [tool_use id=t1 get_weather{city: Tokyo}] [tool_use id=t2 get_weather{city: Paris}]
user:       [tool_result t1: {"temp_c": 15, ...}]     [tool_result t2: {"temp_c": 22, ...}]
assistant:  Tokyo is 15 ยฐC and overcast; Paris is 22 ยฐC and sunny.      โ† stop_reason end_turn

Each step adds the model's output, the tool calls and the tool results to the context, and the whole history is re-sent (and re-billed) on every call. A 20-step episode whose context grows by 3K tokens per step processes about 3K ร— (1 + 2 + โ€ฆ + 20) โ‰ˆ 630K input tokens in total, not 60K. Prompt caching makes the repeated prefix cheap (Prompt Caching & Cost), which is why you never edit earlier turns.

Running tool calls in parallel

When the model emits several independent calls, execute them concurrently - latency becomes the slowest call rather than the sum:

import asyncio

async def execute_all(calls):
    async def one(c):
        try:
            return c.id, await run_tool_async(c.name, c.input), False
        except Exception as e:
            return c.id, f"Error: {e}", True
    return await asyncio.gather(*(one(c) for c in calls))

Only parallelise calls that are independent and read-only, or whose side effects commute. Both APIs let you switch parallel calls off (parallel_tool_calls=False on OpenAI; disable_parallel_tool_use in Claude's tool_choice) when order matters.


Bounds and Termination

A loop without bounds will eventually burn money in a cycle. Every loop needs a success condition (the model answers) and escape hatches:

BoundTypical defaultOn breach
Model calls per episode10-50 (task-dependent)Return partial result with status="max_steps"
Tool calls per episode2-3ร— model callsSame
Identical calls (same tool + same arguments)2Return an error result telling the model to use the earlier result or change approach
Tokens or cost per episodeFrom your p95 on a test set, with headroomStop and escalate
Wall-clock timePer product SLAStop; resume later if the loop is durable

Return a structured result, not an exception: status, final or partial answer, steps, tool calls, tokens, and the trace. That record is what your evals and dashboards consume (Agent Evaluation Basics).


Common Loop Failures

FailureSymptomFix in the loop
Claimed actionFinal answer says "I've refunded it" but no write tool was calledGrade on environment state, not text; in high-stakes flows, generate the confirmation from tool results
RepetitionSame call with the same arguments again and againRepeated-call detection; informative "no results" messages
Hallucinated argumentsIds or values that appear nowhere in the contextSchema validation; lookup tools; "never guess ids" in descriptions
Orphaned tool callsAPI error on the next requestAnswer every call, including on exceptions and timeouts
Context exhaustionForgets the goal, contradicts earlier findingsTruncate large tool outputs; compact history; keep a task summary (Agent Memory)
Premature stopAnswers before finishing all parts of the requestInstructions defining "done"; a verification step before accepting the answer

The first failure is the one the lab catches most often on a small model: the reply describes an action the agent never took, which only an environment-state check detects.


Check Yourself

Check yourself
0 / 5 answered
  1. A Claude response has stop_reason 'max_tokens' and contains a tool_use block. What should the loop do?
  2. What must you do with 'pause_turn'?
  3. A reasoning model returns reasoning items alongside a function_call in the Responses API. What happens to them on the next request?
  4. An episode has 20 model calls and the context grows by about 2,000 tokens per step from a 4,000-token start. Roughly how many input tokens are processed in total?
  5. Why is 'grade on environment state' the fix for claimed actions?

Exercises

Exercise - Break the invariants

Using the lab's loop (or your own), deliberately (a) skip one tool result when the model makes two calls, (b) send two results in two separate messages, (c) drop the assistant turn's tool_calls before appending results. Record the error or behaviour you get from your server or API for each.

Solution

Typical outcomes: (a) an API validation error that a tool call id has no response (OpenAI and Claude both reject this); local servers may accept it but the model then re-issues the call or hallucinates the missing result. (b) Claude requires tool_result blocks at the start of the next user message and rejects interleaving; Chat Completions accepts consecutive tool messages but not other messages in between. (c) The results reference ids that no longer exist - an error on hosted APIs, confusion on local servers. The lesson: history integrity is part of the contract, not a style choice.

Exercise - Add a cost bound

Extend the lab's run_agent with a max_cost_usd bound computed from usage and per-token prices you pass in. When exceeded, stop with status max_cost and return the partial trace. Test it by setting a bound lower than one episode's cost.

Solution

After each call, cost += prompt_tokens ร— p_in + completion_tokens ร— p_out (use cached-token counts where the API reports them). Check the bound before the next model call, not after tool execution, so you never pay for a call you'll discard. Return Result(status="max_cost", ...) with the trace so far.

Exercise - Parallel versus sequential

Add a tool with an artificial 1-second delay and ask a question that needs it three times independently. Measure episode wall time with sequential execution and with asyncio.gather.

Solution

Sequential โ‰ˆ 3 s of tool time plus model time; parallel โ‰ˆ 1 s plus model time. The model-call count is unchanged if the model emits the three calls in one turn; if it emits them one per turn, no executor-side parallelism helps - encourage parallel calls in the instructions.

Study Notes

  • The loop: call model โ†’ execute every tool call โ†’ append assistant turn + all results โ†’ repeat; stop on a final answer or a bound
  • The model is stateless; the message list is its working memory and is re-sent (and re-billed) every call
  • Invariants: one result per call id; all results in one turn; assistant turns (with reasoning) appended unmodified; never edit earlier turns
  • Claude stop reasons: tool_use, end_turn, max_tokens, pause_turn, refusal, model_context_window_exceeded, stop_sequence
  • OpenAI Responses: function_call / function_call_output matched by call_id; pass reasoning items back
  • Return errors as results; bound steps, tool calls, repeats, tokens/cost and time; return a structured result
  • Parallelise independent calls; grade agents on environment state, not on what they say they did

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท