Contents
Map

12 ยท RAG

Agentic & Deep-Research RAG

View as:

Agentic & Deep-Research RAG

Agentic RAG turns retrieval into a tool that a model calls in a loop - deciding what to search for, reading results, and searching again until it can answer - and deep-research systems scale that loop to dozens or hundreds of searches, often across several parallel agents, to produce cited reports.

Learning objectives 50 min
By the end of this page you will be able to:
  • Contrast pipeline RAG with agentic RAG on control flow, latency, cost and failure modes
  • Build an agentic RAG loop with retrieval exposed as a tool, with budgets and stop conditions
  • Describe how deep-research systems are structured (planning, parallel sub-agents, synthesis, citation) and how search agents are trained with RL
  • Evaluate agentic retrieval with trajectory-level metrics and multi-hop benchmarks

From Pipeline to Loop

A pipeline retrieves once with the user's words. Many questions can't be answered that way: "Which of our suppliers were affected by the regulation that came into force after last year's audit?" needs the audit date, then the regulation, then the suppliers - each search depends on the last.

flowchart LR
    Q["โ“ Question"] --> P["๐Ÿง  Model plans<br/>next search"]
    P --> T["๐Ÿ”ง search_docs / web_search / sql"]
    T --> O["๐Ÿ“ฅ Results"]
    O --> J{"โœ”๏ธ Enough evidence?<br/>budget left?"}
    J -->|"no"| P
    J -->|"yes"| A(["๐Ÿ’ฌ Cited answer"])

    style P fill:#d8dfe8,stroke:#b0bac8
    style T fill:#e8e0d4,stroke:#c8b89a
    style J fill:#ddd8e4,stroke:#b8b0c8
    style A fill:#dde4dc,stroke:#b0c4b0
Pipeline RAGAgentic RAG
Control flowFixed: retrieve โ†’ generateModel decides what, when and how often to retrieve
QueriesThe user's (possibly rewritten)Model-written, one per information need
Multi-hopWeakNatural
LatencyPredictable, ~1-3 sVariable, seconds to minutes
CostOne retrieval, one generationSeveral to hundreds of calls
TestingDeterministic stagesTrajectories vary run to run

The ideas behind it: interleaving retrieval with chain-of-thought (IRCoT, Trivedi et al., 2023), ReAct's thought-action-observation loop, and self-ask decomposition. Modern tool-calling models do this natively.


Building It

Expose retrieval as a tool with a precise description, and let the agent loop run:

from langchain.agents import create_agent
from langchain_core.tools import tool

@tool
def search_docs(query: str, k: int = 5) -> str:
    """Search the internal knowledge base (policies, runbooks, contracts).
    Use specific keywords and names; call again with a refined query if results are off-topic.
    Returns passages as [id] source: text."""
    hits = hybrid_search(query, k=k)                     # your retriever + reranker
    return "\n".join(f"[{h.id}] {h.source}: {h.text}" for h in hits)

agent = create_agent(
    model="anthropic:<model>",                           # any tool-calling chat model
    tools=[search_docs, run_sql],
    system_prompt=(
        "Answer questions about the company using the tools. Search as many times as you need, "
        "one information need per query. Cite passage ids like [id] for every factual claim. "
        "Stop when you have enough evidence or after 8 searches; say what you couldn't find."
    ),
)
result = agent.invoke({"messages": [{"role": "user", "content": question}]})

Design points that decide whether it works:

  • Tool descriptions are prompts. Say what the corpus contains, how to phrase queries, and what comes back. Return passage ids and sources so the final answer can cite.
  • Budgets and stop conditions. Cap searches, tokens and wall-clock time; tell the model the cap. Without them, agents over-search on hard questions and loop on unanswerable ones.
  • Parallel calls. Independent sub-questions can be searched in one turn with parallel tool calls.
  • Context growth. Each tool result stays in the context. Return concise passages, and clear or summarise old results for long sessions (see Context Engineering).
  • Untrusted content. Web pages and documents can carry prompt injections that now steer an agent with tools - see Prompts in Production.
  • Hybrid routing. Send simple questions down a fast pipeline and only escalate to the agent when the question is multi-part or the pipeline's retrieval confidence is low.

The agent loop itself - message handling, tool-result formatting, error handling - is covered in The Agent Loop.


Deep Research Systems

Deep-research products (Google's Gemini Deep Research, OpenAI's deep research, Anthropic's Research feature, Perplexity and open-source equivalents) run long agentic retrieval sessions - many searches over the web and internal sources, reading full pages, and writing a report with citations.

flowchart TD
    U["๐Ÿง‘ Research question"] --> L["๐Ÿงญ Lead agent<br/>clarify, plan, split into sub-questions"]
    L --> S1["๐Ÿ”Ž Sub-agent 1<br/>own context, own searches"]
    L --> S2["๐Ÿ”Ž Sub-agent 2"]
    L --> S3["๐Ÿ”Ž Sub-agent 3"]
    S1 & S2 & S3 --> L2["๐Ÿงญ Lead agent<br/>merge findings, find gaps,<br/>spawn more if needed"]
    L2 --> W["โœ๏ธ Write report"]
    W --> C["๐Ÿ”— Citation pass<br/>attach and verify sources"]
    C --> R(["๐Ÿ“„ Cited report"])

    style L fill:#d8dfe8,stroke:#b0bac8
    style S1 fill:#e8e2d9,stroke:#ccc4b8
    style S2 fill:#e8e2d9,stroke:#ccc4b8
    style S3 fill:#e8e2d9,stroke:#ccc4b8
    style C fill:#dde4dc,stroke:#b0c4b0

Anthropic's description of its multi-agent research system gives useful numbers:

  • An orchestrator with parallel sub-agents outperformed a single agent by 90.2% on their internal research evaluation.
  • Token usage alone explained about 80% of the variance in performance on the BrowseComp benchmark. Multi-agent research works largely by spending more tokens in parallel, isolated contexts.
  • Agents used about 4ร— the tokens of a chat interaction, and multi-agent systems about 15ร—. That is only worth it for valuable, broad, parallelisable questions.
  • Their lessons: teach the orchestrator how to delegate (clear objectives and boundaries for each sub-agent), scale effort to query complexity, start searches broad and then narrow, and evaluate end states rather than exact paths.

Training search agents

Reasoning models are now trained to search inside their reasoning with reinforcement learning. Search-R1 (Jin et al., 2025) trains a model to interleave <search> calls with reasoning, rewarded only on the final answer, and improves multi-hop QA substantially over prompting-based RAG. Frontier reasoning models are trained with tool use in the loop in similar ways, which is why they search more deliberately than prompted pipelines. The RL machinery is covered in RL for LLMs.


Evaluating Agentic Retrieval

LevelMetrics
Final answerCorrectness vs reference; groundedness; citation recall and precision
TrajectoryNumber of searches and tokens; redundant queries; whether the needed documents were ever retrieved (retrieval recall over the whole trajectory)
EfficiencyLatency, cost per question, success at a fixed budget
RobustnessBehaviour on unanswerable questions (does it stop?), injected content, tool errors

Benchmarks: multi-hop QA (HotpotQA, MuSiQue, 2WikiMultiHopQA), FRAMES (multi-document reasoning with retrieval), BrowseComp (1,266 hard-to-find facts on the open web), and GAIA (general assistant tasks with tools). Because trajectories vary, run each item several times and report mean and variance. Agent-level evaluation is developed further in Production Agents.


Check Yourself

Check yourself
0 / 4 answered
  1. Which question most justifies agentic RAG over a pipeline?
  2. According to Anthropic's analysis of its research system, what explained most of the variance in BrowseComp performance?
  3. Your agent sometimes performs 40 searches on unanswerable questions. What are the two most direct fixes?
  4. Why do deep-research systems use sub-agents with separate contexts rather than one long-running agent?

Exercises

Exercise - Pipeline vs agent on multi-hop questions

Build 30 two-hop questions over a corpus you control (e.g. "Who manages the team that owns service X?" where ownership and management are in different documents). Compare a hybrid pipeline (retrieve 8, answer) with an agent that has the same retriever as a tool and a budget of 6 searches. Report accuracy, mean searches, tokens and p50 latency.

Solution

Expect the agent to win clearly on two-hop questions (the pipeline rarely retrieves both hops with one query) at several times the tokens and latency. On single-hop questions added as a control, the pipeline should match the agent at a fraction of the cost - the argument for routing.

Exercise - Design a research agent's budget

Specify the budgets, stop conditions, sub-agent instructions and citation checks for an internal "competitive analysis" research agent that must finish within 5 minutes and $2 per report.

Solution

Example: lead agent plans at most 5 sub-questions; each sub-agent gets 10 searches and 40K tokens, must return findings with source URLs and a confidence note; lead agent may spawn one follow-up round for gaps; a final citation pass checks every claim against its source and drops unsupported ones; hard stop at 4.5 minutes with a partial-report fallback. Budget checks use token counts from each call.

Study Notes

Must-know:

  • Agentic RAG: retrieval as a tool in a model-driven loop; natural multi-hop; variable latency and cost
  • Tool descriptions, budgets, stop conditions, parallel calls, context hygiene and injection defence decide success
  • Route easy questions to a pipeline, hard ones to the agent
  • Deep research: lead agent + parallel sub-agents in isolated contexts + synthesis + citation pass; performance tracks tokens spent
  • Search agents are increasingly trained with RL (Search-R1) rather than only prompted
  • Evaluate final answers, trajectories, efficiency and robustness; run items multiple times

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท