Contents
Map

12 ยท RAG

Long Context vs RAG vs CAG

View as:

Long Context vs RAG vs CAG

With context windows of hundreds of thousands to millions of tokens, you can often skip retrieval and put the whole corpus in the prompt - either on every call (long-context prompting) or once, with its KV cache reused across queries (cache-augmented generation, CAG). This chapter is the decision guide.

Learning objectives 40 min
By the end of this page you will be able to:
  • Describe long-context prompting, RAG and cache-augmented generation and what each costs per query
  • Summarise the evidence comparing long context and RAG on quality and cost
  • Compute the per-query cost and, for self-hosting, the KV-cache memory of putting a corpus in context
  • Choose an approach - or a hybrid such as routing or retrieve-then-read-long - for a given corpus and workload

Three Ways to Put Knowledge in Context

flowchart TD
    subgraph RAG["๐Ÿ“ก RAG"]
        R1["Retrieve top-k chunks<br/>per query"] --> R2["Small context<br/>(~2-20K tokens)"]
    end
    subgraph LC["๐Ÿ“œ Long-context prompting"]
        L1["Send the whole corpus<br/>with every query"] --> L2["Huge context<br/>(100K-1M+ tokens)"]
    end
    subgraph CAG["โšก Cache-augmented generation"]
        C1["Prefill the corpus once;<br/>keep its KV cache"] --> C2["Each query reuses the cache;<br/>only the question is new"]
    end

    style RAG fill:#dde4dc,stroke:#b0c4b0
    style LC fill:#e8e0d4,stroke:#c8b89a
    style CAG fill:#d8dfe8,stroke:#b0bac8
  • RAG pays for retrieval infrastructure and risks missing evidence, but each call is small and cheap and the corpus can be any size.
  • Long-context prompting gives the model everything - no retrieval misses, and whole-document reasoning - but every call pays to process the full corpus, and quality degrades as context grows (see Context Engineering).
  • CAG (Chan et al., 2024) keeps long context's completeness while avoiding repeated prefill: the corpus's KV cache is computed once and reused. With hosted APIs, CAG is simply prompt caching of a corpus-sized prefix; self-hosted, it means keeping (or offloading and reloading) a large KV cache.

What the Evidence Says

StudyFinding
Xu et al. (ICLR 2024)Retrieval helped even long-context models; a 4K-context model with retrieval matched a 16K long-context fine-tune on long-document QA, at lower cost
Li et al. (EMNLP 2024 Industry)With sufficient resources, long-context models outperformed RAG on average across long-context benchmarks, while RAG was far cheaper. Their Self-Route method - try RAG first and let the model declare whether the retrieved chunks suffice, falling back to long context only when they don't - kept quality close to long context at a fraction of the cost
Chan et al. (2024)CAG matched or beat RAG on HotpotQA and SQuAD when the whole knowledge source fit in context, and removed retrieval latency
RULER, NoLiMa, Context Rot (2024-25)Effective context is well below advertised windows for hard tasks, and reliability declines with length

The consistent picture: long context wins on completeness when the corpus fits and the task needs broad reading; RAG wins on cost and scale; hybrids capture most of both.


The Cost Arithmetic

Example: a 400K-token corpus, 20,000 questions per day. Prices: $3/M input, $0.30/M cached input, $15/M output; 500 output tokens per answer.

ApproachInput per queryCost per queryPer day
Long context, uncached400K ร— $3/M = $1.20โ‰ˆ $1.21โ‰ˆ $24,150
Long context with prompt caching (CAG, warm cache)400K ร— $0.30/M = $0.12โ‰ˆ $0.13โ‰ˆ $2,550
RAG with 8K tokens of context8K ร— $3/M = $0.024โ‰ˆ $0.032โ‰ˆ $630

Caching cuts long-context cost by about 10ร—, and RAG is still about 4ร— cheaper here - plus faster (prefill for 8K tokens vs reading a 400K cache) and not limited by corpus size. But if the corpus were 40K tokens, cached long context would cost about $0.02 per query - the same as RAG, with no retrieval to build or maintain.

Self-hosted CAG: KV memory

KV-cache size per token = 2 (K and V) ร— layers ร— KV heads ร— head dimension ร— bytes per value. For a model with 64 layers, 8 KV heads (GQA), head dimension 128, in bf16:

2 ร— 64 ร— 8 ร— 128 ร— 2 bytes = 262,144 bytes โ‰ˆ 256 KB per token
400K-token corpus โ†’ ~105 GB of KV cache

That is more than one H100's memory for a single cached corpus - before the model weights. Self-hosted CAG needs FP8 KV caches, KV offloading to CPU or SSD (e.g. LMCache), or a smaller corpus (see The Modern Serving Stack).


Decision Guide

SituationRecommendation
Corpus fits comfortably (under about 10-20% of the window), stable, many queriesCached long context (CAG) - simplest, no retrieval misses
Corpus fits, but questions are narrow lookups at high volumeRAG or Self-Route - cheaper and faster per query
Corpus far larger than the window, or growingRAG (optionally agentic)
Questions need whole-document reasoning (a contract, a codebase module)Retrieve documents, then read them whole: retrieve at document level and pass complete documents, not chunks
Per-user or per-tenant document setsRAG with access filters - a shared cached prefix can't vary by user
Content changes hourlyRAG - a changed corpus invalidates the cache
Strict citation requirementsEither, with quote extraction or API-native citations; RAG gives passage-level provenance for free
flowchart TD
    A{"Does the corpus fit well<br/>within the window?"} -->|"no"| RAG1["๐Ÿ“ก RAG / agentic RAG"]
    A -->|"yes"| B{"Same corpus for all users<br/>and changes rarely?"}
    B -->|"no"| RAG2["๐Ÿ“ก RAG with filters"]
    B -->|"yes"| C{"Many queries per<br/>cache lifetime?"}
    C -->|"yes"| CAGN["โšก Cached long context (CAG)"]
    C -->|"no"| D{"Needs broad reading<br/>or whole-document reasoning?"}
    D -->|"yes"| LCN["๐Ÿ“œ Long context"]
    D -->|"no"| RAG3["๐Ÿ“ก RAG or Self-Route"]

    style CAGN fill:#d8dfe8,stroke:#b0bac8
    style LCN fill:#e8e0d4,stroke:#c8b89a
    style RAG1 fill:#dde4dc,stroke:#b0c4b0
    style RAG2 fill:#dde4dc,stroke:#b0c4b0
    style RAG3 fill:#dde4dc,stroke:#b0c4b0

Whatever you choose, measure on your questions: build the eval once (RAG Evaluation) and run both approaches at your real corpus size.


Check Yourself

Check yourself
0 / 4 answered
  1. With hosted APIs, what is cache-augmented generation in practice?
  2. What did Li et al.'s Self-Route do?
  3. Why is CAG a poor fit for a multi-tenant assistant where each customer sees only their own documents?
  4. A 64-layer model with 8 KV heads of dimension 128 caches in bf16. Roughly how much KV memory does a 100K-token cached corpus need?

Exercises

Exercise - Price your own case

Pick a real corpus (or assume one) and fill in: corpus tokens, queries per day, cache hit rate, prices for your provider. Compute daily cost for uncached long context, cached long context and RAG with your expected context size. At what corpus size do cached long context and RAG cost the same per query?

Solution

Break-even (ignoring output): corpus_tokens ร— p_cached = rag_tokens ร— p_input. With p_cached = 0.1 ร— p_input, break-even corpus size โ‰ˆ 10 ร— RAG context size - e.g. an 8K-token RAG context breaks even with an ~80K-token cached corpus. Below that, cached long context is not more expensive than RAG and avoids retrieval misses.

Exercise - Compare quality

For a 50-100K-token document set that fits in context, write 40 questions (half narrow lookups, half requiring synthesis across documents). Compare RAG (top 8 chunks) with full long context on accuracy and cost.

Solution

Expect similar accuracy on narrow lookups and a long-context advantage on synthesis questions, with RAG far cheaper unless the corpus is cached. The split by question type is the practical output - it tells you whether routing (Self-Route style) is worthwhile.

Study Notes

Must-know:

  • RAG: small cheap calls, any corpus size, risk of retrieval misses
  • Long context: completeness and whole-document reasoning; cost scales with corpus; quality degrades with length
  • CAG: precompute and reuse the corpus's KV cache = prompt caching for hosted APIs
  • Evidence: long context often wins on quality when it fits; RAG wins on cost; Self-Route hybrids get most of both
  • Cost: corpus ร— cached price vs RAG context ร— input price; break-even corpus โ‰ˆ 10ร— the RAG context at a 0.1ร— cache price
  • KV memory = 2 ร— layers ร— KV heads ร— head dim ร— bytes per token - large corpora need FP8 KV or offloading
  • Per-tenant data, frequent changes and huge corpora point to RAG

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท