Contents
Map

12 ยท RAG

RAG Fundamentals

View as:

RAG Fundamentals

Retrieval-augmented generation (RAG) answers a question by first retrieving relevant passages from a corpus you control and then having a language model generate an answer grounded in, and citing, those passages.

Learning objectives 45 min
By the end of this page you will be able to:
  • Draw the indexing and query pipelines of a RAG system and name the design decision at each stage
  • Explain which LLM limitations RAG addresses and which it does not
  • Choose between RAG, fine-tuning and long-context prompting for a knowledge problem
  • Map the classic failure modes of naive RAG to the chapters that fix them

The Idea

Lewis et al. (2020) combined a parametric memory - knowledge stored in model weights - with a non-parametric memory: a dense index over Wikipedia that the model could retrieve from at generation time. The idea became the dominant way to put private, fresh or verifiable knowledge in front of a model: keep facts in a store you can update and audit, and put the relevant ones in the context window when needed.

flowchart TD
    subgraph IDX["๐Ÿ“ฅ Indexing (offline, incremental)"]
        D["๐Ÿ“„ Sources<br/>PDFs, wikis, tickets, DBs"] --> P["๐Ÿงน Parse & clean<br/>layout, tables, metadata"]
        P --> CH["โœ‚๏ธ Chunk"] --> EMB["๐Ÿงฌ Embed<br/>(+ sparse index)"] --> VS[("๐Ÿ—„๏ธ Index<br/>vectors + text + metadata")]
    end
    subgraph QRY["๐Ÿ” Query (online, per request)"]
        Q["โ“ Question"] --> QT["โœ๏ธ Query rewrite<br/>(optional)"]
        QT --> RET["๐Ÿ“ก Retrieve<br/>dense / sparse / hybrid<br/>+ metadata filters"]
        RET --> RR["๐ŸŽฏ Rerank"]
        RR --> CTX["๐Ÿงฉ Assemble context<br/>ordered, tagged, cited"]
        CTX --> GEN["๐Ÿค– Generate<br/>grounded answer + citations"]
    end
    VS -.-> RET

    style IDX fill:#e8e2d9,stroke:#ccc4b8
    style QRY fill:#d8dfe8,stroke:#b0bac8
    style GEN fill:#dde4dc,stroke:#b0c4b0
StageKey decisionChapter
Parse and chunkHow documents become retrievable unitsDocument Processing & Chunking
Embed and indexWhich embedding model, which index, how much compressionEmbeddings & Vector Search
Retrieve and rerankDense, sparse or hybrid; which reranker; how many candidatesRetrieval & Reranking
GenerateHow context is arranged, how the model cites and abstainsGrounded Generation & Citations
EvaluateRetrieval metrics, groundedness, answer qualityRAG Evaluation

One invariant to remember: the query must be embedded with the same model (and version) as the corpus. Vectors from different models live in different spaces; mixing them returns plausible-looking but meaningless results with no error. Changing the embedding model means re-embedding the whole corpus.


A Minimal Pipeline

Frameworks such as LangChain, LlamaIndex and Haystack package these steps, but the core fits in a few lines:

import numpy as np
from sentence_transformers import SentenceTransformer
from langchain_text_splitters import RecursiveCharacterTextSplitter

documents = {"refund-policy.md": "...", "sla.md": "..."}

# Index (offline)
splitter = RecursiveCharacterTextSplitter(chunk_size=800, chunk_overlap=100)
chunks = [(src, text) for src, doc in documents.items() for text in splitter.split_text(doc)]
embedder = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
chunk_vecs = embedder.encode([t for _, t in chunks], normalize_embeddings=True)

# Query (online)
def retrieve(question: str, k: int = 5):
    q = embedder.encode([question], normalize_embeddings=True)[0]
    scores = chunk_vecs @ q                       # cosine similarity: vectors are normalised
    return [(chunks[i][0], chunks[i][1]) for i in np.argsort(-scores)[:k]]

def build_prompt(question: str, hits) -> str:
    context = "\n".join(f'<doc id="{i}" source="{src}">{text}</doc>' for i, (src, text) in enumerate(hits, 1))
    return ("Answer using only the documents below and cite them as [id]. "
            "If they don't contain the answer, say you don't know.\n\n"
            f"{context}\n\nQuestion: {question}")

answer = llm(build_prompt(question, retrieve(question)))   # any chat model

A production system adds document parsing, metadata and access filters, hybrid retrieval, a reranker, citation checking, caching, evaluation and monitoring - the rest of this module. The code lab measures how much each retrieval choice actually matters.


What RAG Solves - and What It Doesn't

LLM limitationHow RAG helpsRemaining risk
Knowledge cutoffThe index can be updated minutes after a document changesStale index if ingestion lags or fails
Private knowledgeYour documents are retrieved at query time; nothing is trained into weightsAccess control must be enforced at retrieval
HallucinationThe answer can be grounded in retrieved textThe model can still ignore, misread or embellish the context
VerifiabilityAnswers cite sources a user can checkCitations can point to passages that don't support the claim

RAG is the wrong tool - or needs help - for:

  • Aggregations and exact lookups over structured data ("total revenue by region last quarter"): use SQL or an API as a tool.
  • Computation: use code execution.
  • Questions needing synthesis across a whole corpus ("what are the main themes in 10,000 reports?"): top-k retrieval sees a few chunks; graph-based or map-reduce approaches (see Advanced RAG Patterns) or agentic research (see Agentic & Deep-Research RAG) fit better.
  • Small, stable corpora that fit in the context window: just put them in the prompt and cache it (see Long Context vs RAG vs CAG).

RAG, Fine-Tuning or Long Context?

NeedRAGFine-tuningLong context (+ caching)
Facts that change weeklyโœ… Update the indexโŒ Retrain each timeโœ… If the corpus fits
Citations and audit trailโœ… Source per chunkโŒโœ… With quote extraction
Corpus of millions of documentsโœ…โŒโŒ
New output style, format or domain behaviourโŒโœ…โš ๏ธ Via examples
Per-user access controlโœ… Filter at retrievalโŒโš ๏ธ Per-user prompts
Lowest latency per queryโš ๏ธ Retrieval adds ~50-300 msโœ…โŒ Long prefill unless cached

Fine-tuning changes behaviour reliably and injects facts unreliably - fine-tuned models still hallucinate about their training data, and updating facts means retraining. The common production combination is a fine-tuned or well-prompted model for behaviour plus RAG for knowledge. The decision between RAG and long-context prompting is covered in Long Context vs RAG vs CAG.


Why Naive RAG Fails

"Naive RAG" - fixed-size chunks, one embedding model, top-k, stuff the prompt - is a fine prototype and a poor product. Its failure modes, and where each is fixed:

SymptomRoot causeFix (chapter)
Right document exists but isn't retrievedVocabulary mismatch, rare terms, poor chunksHybrid retrieval, contextual chunks, query rewriting (03, 04)
Retrieved but ranked below noiseBi-encoder imprecisionReranking (04)
Answer spans two chunksChunk boundariesStructure-aware chunking, parent-child retrieval (03)
Tables and figures missingParsing loses layoutLayout-aware parsing, multimodal retrieval (03, 07)
Context present but answer wrong or embellishedGeneration ignores or misuses contextContext ordering, quote-then-answer, citation checks (05)
Multi-hop questions failOne retrieval can't gather all hopsIterative or agentic retrieval (08)
Silent quality driftNo measurementRetrieval and groundedness evals in CI and production (06, 11)

Check Yourself

Check yourself
0 / 4 answered
  1. After upgrading the embedding model used for queries, every answer becomes irrelevant but no errors appear. What happened?
  2. Which question is RAG over documents least suited to?
  3. A 40-page product manual changes monthly and every answer must cite the page. Which approach is simplest?
  4. Why does fine-tuning not replace RAG for fast-changing facts?

Exercises

Exercise - Run and break the minimal pipeline

Run the minimal pipeline above on 20 documents of your choice (e.g. a product's docs pages). Write 10 questions with known answers. Record how many retrieve the right chunk in the top 3. Then deliberately break it three ways - embed queries with a different model, set chunk_size to 50, set k to 1 - and record the effect of each.

Solution

A different query model collapses retrieval to near random with no error; tiny chunks lose context and hurt answerability even when retrieval "hits"; k=1 loses recall on questions whose answer is spread across chunks. The exercise shows why each parameter needs measurement.

Exercise - Choose an approach

For each case choose RAG, fine-tuning, long context, a tool (SQL/API) or a combination, and justify it in one sentence: (a) a support bot over 50,000 help-centre articles; (b) a model that must write discharge summaries in a hospital's house style; (c) "how many open P1 incidents do we have?"; (d) Q&A over a single 80-page contract.

Solution

(a) RAG - large, changing corpus with citations. (b) Fine-tuning (or strong examples) for style, plus RAG for patient-specific facts. (c) A tool calling the incident system's API - an exact, live count. (d) Long context with caching and quote-then-answer - it fits in the window and needs whole-document reasoning.

Study Notes

Must-know:

  • RAG = retrieve relevant passages from your corpus, then generate a cited answer grounded in them (Lewis et al., 2020)
  • Two pipelines: indexing (parse, chunk, embed, index) and query (rewrite, retrieve, rerank, assemble, generate)
  • Same embedding model for queries and corpus - mismatches fail silently
  • RAG helps with freshness, private data, grounding and citations; not with aggregations, computation or whole-corpus synthesis
  • Fine-tune for behaviour, retrieve for knowledge; small stable corpora can simply go in a cached prompt
  • Naive RAG's failures map to chunking, hybrid retrieval, reranking, grounded generation, agentic retrieval and evaluation

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท