Contents
Map

12 · RAG

Advanced RAG Patterns

View as:

Advanced RAG Patterns

Advanced RAG patterns change when retrieval happens (adaptive and corrective RAG), what is indexed (knowledge graphs, page images, tables) and how the pipeline is composed (routing across modules) - each targeting a failure that single-shot vector retrieval can't fix.

Learning objectives 60 min
By the end of this page you will be able to:
  • Place a RAG design on the axes of control flow (fixed vs adaptive) and index type (text, graph, multimodal, structured)
  • Explain Self-RAG, CRAG, FLARE, Adaptive-RAG and Speculative RAG and what each needs to work
  • Describe GraphRAG's indexing pipeline and local/global search, and compare it with LightRAG and LazyGraphRAG on cost
  • Choose between caption-based, CLIP-style and ColPali-style retrieval for visually rich documents

A Map, Not a Ladder

Surveys describe an evolution from naive to advanced (pre- and post-retrieval improvements) to modular RAG (swappable, routed components). In practice these are independent design choices you combine:

mindmap
  root((RAG design))
    Control flow
      Fixed pipeline
      Routed by query type
      Adaptive retrieval
        Self-RAG
        FLARE
        Adaptive-RAG
      Corrective
        CRAG
      Agentic loop
    Index
      Text chunks
        dense
        sparse
        late interaction
      Knowledge graph
        GraphRAG
        LightRAG
      Page images
        ColPali
      Structured data
        SQL and APIs
    Generation
      Single pass
      Draft and verify
        Speculative RAG
      Map reduce
AxisChoiceSolvesCost
Pre-/post-retrievalQuery rewriting, hybrid, reranking, compressionRecall and precision of single-shot retrievalLow - the "advanced RAG" baseline (04)
RoutingPick a retriever or source per queryMixed query types (docs, SQL, web)A router to build and evaluate
Adaptive / correctiveDecide whether, when and how often to retrieveUnnecessary retrieval; bad retrievalExtra model calls; variable latency
AgenticThe model plans and issues searches in a loopMulti-hop, research tasksHighest latency and cost (08)
Graph indexEntities and relations extracted into a graphRelationship and whole-corpus questionsExpensive indexing
Multimodal indexPage images or image embeddingsCharts, tables, scans, slidesStorage, vision models

Routing (Modular RAG)

A router classifies each query and sends it to the right module: vector search over docs, SQL over a warehouse, a web search, a graph query, or no retrieval at all.

flowchart LR
    Q["❓ Query"] --> R{"🧭 Router<br/>(classifier or LLM<br/>with structured output)"}
    R -->|"policy / how-to"| V["📚 Hybrid doc retrieval"]
    R -->|"numbers / aggregates"| S["🗃️ Text-to-SQL"]
    R -->|"current events"| W["🌐 Web search"]
    R -->|"small talk / general"| N["💬 Answer directly"]
    V & S & W & N --> G(["🤖 Grounded answer"])

    style R fill:#d8dfe8,stroke:#b0bac8
    style G fill:#dde4dc,stroke:#b0c4b0

Implement the router with structured output (an enum of routes) or a small classifier, and evaluate it like any classifier - a mis-route is a guaranteed wrong answer. With tool-calling models, routing is often expressed as tools the model chooses between, which is the bridge to agentic RAG.


Adaptive and Corrective Retrieval

PatternMechanismRequirementsUse when
Self-RAG (Asai et al., 2023)A fine-tuned model emits reflection tokens: Retrieve (yes / no / continue), IsRel (is a passage relevant), IsSup (fully / partially / not supported), IsUse (usefulness 1-5); decoding uses them to decide when to retrieve and which output to keepA model trained to emit the tokens; open-source checkpoints existResearch and open-model stacks; the idea (critique retrieval and support) is widely reused with prompted judges
CRAG (Yan et al., 2024)A lightweight retrieval evaluator grades results Correct, Incorrect or Ambiguous; correct results are refined (decomposed, irrelevant strips filtered), incorrect ones replaced by web search, ambiguous ones combinedA retrieval grader (the paper fine-tunes a small T5; an LLM judge also works) and a fallback sourceCorpora with coverage gaps where a fallback source is acceptable
FLARE (Jiang et al., 2023)Generate a tentative next sentence; if it contains low-probability tokens, use it as a query, retrieve, and regenerateToken probabilities from the generatorLong-form generation where information needs emerge while writing
Adaptive-RAG (Jeong et al., 2024)A small classifier predicts query complexity and routes to no retrieval, single-step or multi-step retrievalA labelled complexity classifierMixed traffic where most queries are simple
Speculative RAG (Wang et al., 2024)A small specialist model writes several answer drafts in parallel, each from a different subset of retrieved documents; a larger generalist model verifies and picks the bestTwo models; parallel inferenceMany retrieved documents, latency-sensitive quality gains

A corrective loop is simple to build with any model and a graph library:

from typing import Literal, TypedDict
from langgraph.graph import StateGraph, START, END

class State(TypedDict):
    question: str
    docs: list[str]
    grade: str
    answer: str

def retrieve(s: State):   return {"docs": search_docs(s["question"], k=8)}
def grade(s: State):      return {"grade": grade_retrieval(s["question"], s["docs"])}   # "correct" | "ambiguous" | "incorrect", via structured output
def web(s: State):        return {"docs": (s["docs"] if s["grade"] == "ambiguous" else []) + web_search(s["question"], k=5)}
def generate(s: State):   return {"answer": grounded_answer(s["question"], s["docs"])}

def route(s: State) -> Literal["web", "generate"]:
    return "generate" if s["grade"] == "correct" else "web"

g = StateGraph(State)
for name, fn in [("retrieve", retrieve), ("grade", grade), ("web", web), ("generate", generate)]:
    g.add_node(name, fn)
g.add_edge(START, "retrieve"); g.add_edge("retrieve", "grade")
g.add_conditional_edges("grade", route)
g.add_edge("web", "generate"); g.add_edge("generate", END)
crag = g.compile()

Every extra judge call adds latency and a failure mode of its own; evaluate the grader's precision and recall on labelled retrievals before trusting it.


Graph-Based RAG

Vector retrieval answers "find passages like this question". It struggles with relationship questions ("which suppliers are shared by our two largest customers?") and global questions ("what are the main risk themes across 5,000 incident reports?"), because the answer is not in any single chunk.

GraphRAG

Microsoft's GraphRAG (Edge et al., 2024) builds a knowledge graph and a hierarchy of summaries at index time:

flowchart TD
    D["📚 Corpus"] --> E["🤖 LLM extracts entities,<br/>relationships, claims per chunk"]
    E --> KG["🕸️ Knowledge graph"]
    KG --> C["🧩 Community detection<br/>(hierarchical Leiden)"]
    C --> S["📝 LLM summary per community,<br/>at each level of the hierarchy"]
    S --> L["🔎 Local search<br/>entity neighbourhoods + text"]
    S --> GL["🌍 Global search<br/>map-reduce over community summaries"]

    style E fill:#e8e0d4,stroke:#c8b89a
    style C fill:#d8dfe8,stroke:#b0bac8
    style GL fill:#dde4dc,stroke:#b0c4b0
  • Local search starts from entities mentioned in the question and gathers their relationships, community reports and source text - for entity-centric questions.
  • Global search maps the question over community summaries at a chosen level and reduces the partial answers - for corpus-wide "sensemaking" questions, where the paper reports large gains in comprehensiveness and diversity over vector RAG.
  • DRIFT search combines the two, starting global and refining locally.
pip install graphrag
graphrag init --root ./project          # writes settings.yaml; put documents in ./project/input
graphrag index --root ./project         # LLM extraction + communities + summaries (the expensive step)
graphrag query --root ./project --method global "What are the main themes in these reports?"
graphrag query --root ./project --method local  "Who works with the procurement team?"

Cost is the catch: indexing runs LLM calls over every chunk and every community, typically costing many times the corpus's token count, and re-indexing on changes is expensive. Cheaper variants:

VariantIdeaTrade-off
LazyGraphRAG (Microsoft, 2024)Defers summarisation to query time; indexes with NLP noun-phrase extraction and co-occurrenceMicrosoft reports indexing cost equal to vector RAG and 0.1% of full GraphRAG, with comparable global-query quality
LightRAG (Guo et al., 2024)Graph plus vector index with dual-level (entity and theme) retrieval; supports incremental updatesSimpler and cheaper; less thorough than full community summaries
HippoRAG (Gutiérrez et al., 2024)Personalised PageRank over an entity graph seeded by query entitiesStrong multi-hop recall at low cost

Managed options exist too - e.g. Amazon Bedrock Knowledge Bases supports GraphRAG backed by Neptune Analytics. If your data is already a graph or relational (a CRM, a product catalogue), query it directly (Cypher, GQL, SQL) as a tool rather than extracting a graph from text.


Multimodal and Visually Rich Documents

ApproachHowBest forLimits
Caption / transcribeA vision-language model describes images and converts tables to Markdown at ingest; index the textMost pipelines; reuses text retrievalLoses detail the caption omits; ingest cost
Joint image-text embeddings (CLIP-style)Embed images and text into one spacePhotos, product images, "find a picture of..."Weak on text-heavy pages and charts
Late interaction over page images (ColPali, Faysse et al., 2024)A vision-language model embeds each page image as ~1,000 patch vectors; queries are scored with MaxSim against patches - no OCRSlides, reports, forms, charts - pages where layout carries meaningStorage per page; a VLM is needed to answer from the retrieved page images

ColPali-style models set the state of the art on the ViDoRe visual document retrieval benchmark and remove the fragile parse-OCR-chunk pipeline for visually rich documents. The generator then receives the retrieved page images, so it must be a multimodal model.


Choosing

Question shapePattern
Single-hop factual, mixed phrasingHybrid + rerank (advanced RAG)
Several distinct source typesRouting
Many queries need no retrievalAdaptive-RAG-style classifier
Corpus has gaps; web fallback acceptableCRAG
Relationships between entities across documentsGraphRAG (local) / HippoRAG, or query an existing graph
"Summarise themes across the whole corpus"GraphRAG global / LazyGraphRAG, or map-reduce
Multi-step research, comparisonsAgentic RAG (08)
Slides, scans, chartsColPali-style page retrieval or VLM captioning

Add complexity only when your eval shows the simpler pipeline failing on that question shape.


Check Yourself

Check yourself
0 / 4 answered
  1. What do Self-RAG's reflection tokens require that CRAG does not?
  2. What does Speculative RAG do?
  3. Which question most clearly calls for GraphRAG global search rather than vector RAG?
  4. Why does ColPali not need OCR?

Exercises

Exercise - Cost out GraphRAG

Your corpus is 20 million tokens. Assume GraphRAG indexing processes about 3× the corpus in input tokens and 0.5× in output tokens across extraction and summarisation, with a model priced at $0.50/M input and $2/M output. Estimate the indexing cost, and the cost of re-indexing 10% of the corpus monthly if incremental updates are not available.

Solution

Input: 60M × $0.50/M = $30; output: 10M × $2/M = $20; about $50 per full index with a small, cheap model (several times more with a frontier model). Without incremental updates, a monthly change forces a full re-index: $50/month here - small in dollars, but hours of wall-clock time and quota. The ratios, not these example prices, are the point: indexing cost scales with corpus size and model price, and query-time approaches (LazyGraphRAG) or incremental ones (LightRAG) change the economics.

Exercise - Build a router

Define three routes for a company assistant (docs, sql, direct). Write 60 labelled queries (20 per route, including ambiguous ones) and implement the router with a structured-output enum. Report per-route precision and recall, and decide what should happen on low-confidence routes.

Solution

Typical confusions are numeric questions about policies ("how many days of leave?" - docs, not SQL) and vague requests. A safe default for low confidence is to query docs and SQL in parallel and let the generator use whichever returns evidence, or to ask a clarifying question.

Study Notes

Must-know:

  • RAG patterns are composable choices: control flow (fixed, routed, adaptive, corrective, agentic) × index (text, graph, images, structured)
  • Self-RAG: fine-tuned reflection tokens (Retrieve, IsRel, IsSup, IsUse); CRAG: retrieval grader + refine / web fallback; FLARE: retrieve when a draft sentence is low-confidence; Adaptive-RAG: route by complexity; Speculative RAG: small drafters + large verifier
  • GraphRAG: LLM-extracted graph, Leiden communities, hierarchical summaries; local, global and DRIFT search; expensive indexing
  • LazyGraphRAG (query-time summarisation, ~vector-RAG indexing cost), LightRAG (dual-level, incremental), HippoRAG (PageRank)
  • ColPali: patch-level late interaction over page images, no OCR; needs a multimodal generator
  • Add a pattern only when evals show the simpler pipeline failing on that question shape

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·