Contents
Map

17 ยท Production Agents

Agent Observability

View as:

Agent Observability

Observability for an agent means being able to answer, for any run, what did it do and why - which model calls it made with what context, which tools it called with what arguments and results, how long each step took, what it cost, and how the run ended - and, across runs, how those numbers are trending. The unit is the trace: one tree of spans per run, following the OpenTelemetry GenAI semantic conventions so any backend can read it.

Learning objectives 40 min
By the end of this page you will be able to:
  • Instrument an agent so each run produces a trace with invoke_agent, chat and execute_tool spans and the GenAI attributes that matter
  • Decide what content to capture (prompts, tool arguments, results) and how to handle PII and size
  • Define the dashboards and alerts for an agent in production - outcomes, bounds, cost, latency, tool errors
  • Debug a failed run from its trace, and turn it into an eval task

A Trace per Run

flowchart TD
    R["๐Ÿค– invoke_agent shop_support<br/>run_id, user, outcome, total tokens, cost"]
    R --> C1["๐Ÿง  chat model-x<br/>input/output tokens, cached tokens, finish reason"]
    R --> T1["๐Ÿ› ๏ธ execute_tool find_customer<br/>call id, duration, error?"]
    R --> C2["๐Ÿง  chat model-x"]
    R --> T2["๐Ÿ› ๏ธ execute_tool refund_item"]
    T2 --> AP["๐Ÿ‘ฉโ€โš–๏ธ approval wait<br/>reviewer, decision, wait time"]
    R --> C3["๐Ÿง  chat model-x<br/>final answer"]

    style R fill:#d8dfe8,stroke:#b0bac8
    style C1 fill:#ddd8e4,stroke:#b8b0c8
    style C2 fill:#ddd8e4,stroke:#b8b0c8
    style C3 fill:#ddd8e4,stroke:#b8b0c8
    style T1 fill:#e8e0d4,stroke:#c8b89a
    style T2 fill:#e8e0d4,stroke:#c8b89a
    style AP fill:#dde4dc,stroke:#b0c4b0

The OpenTelemetry GenAI semantic conventions name the operations - chat (a model call), invoke_agent, execute_tool, create_agent, embeddings and others - and the attributes: gen_ai.operation.name, gen_ai.provider.name, gen_ai.request.model, gen_ai.response.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reasons, gen_ai.agent.name, gen_ai.tool.name, gen_ai.tool.call.id, gen_ai.conversation.id. The conventions are still marked in development and evolve release by release, but most frameworks (OpenAI Agents SDK processors, LangChain/LangGraph, ADK, Microsoft Agent Framework, PydanticAI via Logfire) and observability vendors emit or ingest them. Using them means you can change backends without re-instrumenting.

Hand-instrumenting a loop is a few lines:

from opentelemetry import trace

tracer = trace.get_tracer("shop-agent")

def run_agent(request, user):
    with tracer.start_as_current_span("invoke_agent shop_support") as run:
        run.set_attribute("gen_ai.operation.name", "invoke_agent")
        run.set_attribute("gen_ai.agent.name", "shop_support")
        run.set_attribute("app.user.tenant", user.tenant)            # your own attributes, namespaced
        for step in range(MAX_STEPS):
            with tracer.start_as_current_span(f"chat {MODEL}") as span:
                span.set_attribute("gen_ai.operation.name", "chat")
                reply = client.chat.completions.create(model=MODEL, messages=messages, tools=TOOLS)
                span.set_attribute("gen_ai.usage.input_tokens", reply.usage.prompt_tokens)
                span.set_attribute("gen_ai.usage.output_tokens", reply.usage.completion_tokens)
            ...
            for call in reply.choices[0].message.tool_calls or []:
                with tracer.start_as_current_span(f"execute_tool {call.function.name}") as span:
                    span.set_attribute("gen_ai.operation.name", "execute_tool")
                    span.set_attribute("gen_ai.tool.name", call.function.name)
                    span.set_attribute("gen_ai.tool.call.id", call.id)
                    result = dispatch(call)
                    if not result["ok"]:
                        span.set_status(trace.StatusCode.ERROR, result["error"])
        run.set_attribute("app.run.outcome", outcome)                 # done / limit_reached / escalated

Propagate the trace context into tools and sub-agents (HTTP headers for MCP servers and A2A calls) so one run is one trace across services.

Content: Capture, Protect, Bound

Token counts and timings are not enough to debug an agent - you need the prompts, tool arguments and results. The conventions keep content opt-in (in span events or separate log records) because it is large and sensitive. Practical policy:

  • Capture full content in development and for sampled production traffic; always capture it for failed and escalated runs.
  • Redact or tokenise PII and secrets before export; restrict who can read content; set retention to what debugging and audit need.
  • Store large payloads (documents, long tool results) by reference, not inline.
  • Record the prompt version, model version, tool-set version and framework version on every run, so a regression can be tied to a change.

Metrics, Dashboards and Alerts

SignalWhyAlert on
Outcome rate: done / limit reached / escalated / errorThe headline health metricA shift from baseline for a task type
Online eval scores and user feedbackQuality on real trafficSustained drop
Tool error rate per toolBroken integrations, schema drift, bad argumentsSpike on any tool
Steps, tokens and cost per run (p50/p95/p99)Loops, context bloat, runaway fan-outp95 jump; budget breaches
Latency per run and per step (TTFT, tool duration)User experience, provider slownessp95 over the SLO
Approval rate and wait timeRubber-stamping (near 100% approvals) or blocked queuesQueue age
Security signals: blocked policy checks, injection classifier hits, new egress destinationsAttacks in progressAny unusual increase

Most of these are aggregations of span attributes, so the trace is the single source; dashboards are views over it.

Debugging from a Trace

A typical investigation: a customer says the agent confirmed a refund that never happened.

  1. Find the run by conversation or user id; open its trace.
  2. Look at the execute_tool spans: no refund_item span, or a refund_item span with an error status?
  3. Open the last chat span's content: the model saw the error (or never called the tool) and still wrote "refunded" - a claimed action.
  4. Check the versions on the run: did this start with a prompt or model change?
  5. Fix in the right layer (a harness check, templated confirmations - see Lab 13), then add the conversation as a regression task so it stays fixed.

Tools

Tracing backends that understand agent traces include Langfuse and Arize Phoenix (open source, self-hostable), LangSmith, Braintrust, W&B Weave, Pydantic Logfire, and the cloud platforms' own tracing (Google Cloud Trace for Agent Runtime, AWS CloudWatch for Bedrock AgentCore, Azure Monitor for Foundry). Most combine tracing with dataset management and online evaluation. Choose on data residency and self-hosting needs, OpenTelemetry support, and whether the evaluation features fit your workflow - the instrumentation should stay vendor-neutral either way.

Check Yourself

Check yourself
0 / 4 answered
  1. Which OpenTelemetry GenAI operation name marks a tool invocation span?
  2. Why is prompt and tool content opt-in in the GenAI conventions?
  3. Approvals for an agent's refunds run at 99.7% with a median review time of 4 seconds. What does the dashboard suggest?
  4. What should you record on every run so a regression can be traced to a change?

Exercises

Exercise - Instrument Lab 13

Add OpenTelemetry spans to the Lab 13 agent loop following the GenAI conventions, export them to a local Phoenix or Langfuse instance (or the console exporter), and run the full task set. Find a failing run and diagnose it from the trace alone.

Solution

One invoke_agent span per episode with chat and execute_tool children; tokens on chat spans, errors as span status on tool spans, the task id and outcome as attributes. A failed refund_one run shows get_order followed by a final chat span with no refund_item span - the claimed-action pattern - without reading logs.

Exercise - Design the dashboard

For the Lab 14 MCP shop agent in production, specify six dashboard panels and four alerts with thresholds, including one security alert. Justify each threshold.

Solution

Panels: outcome rate by task type; tool error rate by tool; tokens and cost per run (p50/p95); latency per run (p95); approval rate and queue age; guard blocks per hour. Alerts: outcome success drops by more than 3 standard errors from the weekly baseline; any tool's error rate above 5% for 10 minutes; p95 cost per run above twice baseline; guard blocks on write tools above their normal rate (a possible injection campaign - Lab 14's result attack would trip it).

Study Notes

  • One trace per run: invoke_agent -> chat and execute_tool spans; GenAI attributes for model, tokens, finish reasons, agent and tool names, call ids, conversation id
  • Conventions are evolving but widely supported; stay vendor-neutral; propagate context to tools and sub-agents
  • Content capture is opt-in: sample, always keep failures, redact, bound, reference large payloads; record versions
  • Dashboards from span data: outcomes, online evals, tool errors, steps/tokens/cost percentiles, latency, approvals, security signals
  • Debug from the trace, fix in the right layer, add a regression task

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท