Contents
Map

12 ยท RAG

RAG System Design

View as:

RAG System Design

Designing a production RAG system means turning requirements - corpus size, query volume, latency, freshness, access control and quality targets - into concrete choices for ingestion, indexing, retrieval, generation, caching and failure handling, with the numbers to justify them.

Learning objectives 60 min
By the end of this page you will be able to:
  • Run a structured RAG design discussion from requirements to trade-offs
  • Build a per-stage latency budget and identify which stages to optimise
  • Size an index and choose sharding, replication and freshness strategies for 10M-100M+ chunks
  • Design multi-tenancy and access control that fail closed
  • Specify caching layers and graceful degradation, including the risks of semantic caching

A Framework for the Design Discussion

StepQuestions to settle
1. RequirementsCorpus size and growth? Document types? QPS (average, peak)? Latency SLO (p95 TTFT and total)? Freshness (minutes, hours, daily)? Who may see what? Quality bar and cost per query?
2. Two pipelinesIngest (offline, event-driven or batch) and query (online) - draw both
3. ComponentsParsing and chunking; embedding model; index and store; retrieval (hybrid?), reranker; generator; citations
4. NumbersIndex size, latency budget, cost per query, throughput per component
5. Trade-offsLatency vs quality (reranker, query rewriting), cost vs quality (model size, context size), freshness vs complexity
6. OperationsEvaluation gates, monitoring, failure modes, security (RAG in Production)

Reference Architecture

flowchart TD
    subgraph ING["๐Ÿ“ฅ Ingestion (async)"]
        SRC["๐Ÿ—„๏ธ Sources<br/>object storage, wikis, tickets, DBs"] --> EV["๐Ÿ“ฌ Change events / scheduler"]
        EV --> Q1["๐Ÿงพ Work queue"]
        Q1 --> W["โš™๏ธ Workers<br/>parse โ†’ chunk โ†’ enrich โ†’ embed"]
        W --> IDX[("๐Ÿ—„๏ธ Hybrid index<br/>vectors + BM25 + metadata/ACL")]
        W --> DOCS[("๐Ÿ“ฆ Document store<br/>full text, versions")]
    end
    subgraph QRY["๐Ÿ” Query (sync)"]
        U["๐Ÿง‘ User + identity"] --> API["๐Ÿšช API / gateway<br/>auth, rate limits"]
        API --> QP["โœ๏ธ Query processing<br/>rewrite, filters, route"]
        QP --> RET["๐Ÿ“ก Hybrid retrieval<br/>ACL-filtered, top 50"]
        RET --> RR["๐ŸŽฏ Reranker โ†’ top 5-8"]
        RR --> GEN["๐Ÿค– LLM, streaming<br/>grounded + cited"]
        GEN --> OUT(["๐Ÿ’ฌ Answer + citations"])
    end
    IDX -.-> RET
    DOCS -.-> RR
    OUT -.-> FB["๐Ÿ“Š Logs, feedback,<br/>sampled evals"]

    style ING fill:#e8e2d9,stroke:#ccc4b8
    style QRY fill:#d8dfe8,stroke:#b0bac8
    style FB fill:#dde4dc,stroke:#b0c4b0

Latency Budget

A p95 target of about 3 seconds end to end, with streaming so the user sees text in about 1 second:

StageTypical p95Levers
Auth, routing, filters10-30 ms-
Query rewrite (LLM)200-800 msSkip for standalone queries; use a small model
Query embedding10-50 msSelf-host or cache
Hybrid retrieval (ANN + BM25)10-80 msIndex tuning, replicas, filter selectivity
Reranking 50 candidates30-300 msSmaller or GPU-hosted reranker; fewer candidates
LLM time to first token300-1,500 msSmaller model, shorter context, prompt caching, lower reasoning effort
LLM generation1-5 sStream it; concise answers

Rules of thumb: the LLM dominates, so context size, model choice and caching matter more than shaving milliseconds from the vector search; run retrieval branches in parallel; and every optional LLM step (rewrite, grading, reranking with an LLM) needs to earn its latency on the eval.


Scaling the Index

Sizing. 100M documents ร— 1.2 chunks ร— 768 dimensions ร— 4 bytes โ‰ˆ 369 GB of float32 vectors - more than a single machine's RAM for an in-memory HNSW index. Options, from Embeddings & Vector Search: int8 (โ‰ˆ92 GB) or binary with rescoring (โ‰ˆ12 GB in RAM), IVF-PQ or DiskANN, or a distributed managed service.

ConcernApproach
ThroughputReplicate the index; queries are read-only and scale horizontally
SizeShard; each shard searched in parallel, results merged (top-k per shard, then global top-k)
Shard keyHash sharding spreads load evenly but every query fans out to all shards; sharding by tenant or domain lets filtered queries hit one shard
FreshnessEvent-driven upserts: on change, delete the document's chunks by doc_id, then insert new ones; idempotent workers; a dead-letter queue for failures
Big rebuilds (new embedding model, new chunking)Build a new index alongside, evaluate, switch traffic with an alias (blue/green), keep the old one for rollback
Ingestion throughputBatch embedding calls, parallel workers behind a queue, back-pressure on API rate limits

Multi-Tenancy and Access Control

Retrieval must never return a chunk the user isn't allowed to see - once a chunk is in the prompt, the model may repeat it.

Isolation modelHowUse when
Index per tenantSeparate collections/indexesStrict contractual or regulatory isolation; few, large tenants
Namespace / partition per tenantEngine-level partitions in a shared clusterMany tenants of varied size - the common SaaS default
Shared index + mandatory filtertenant_id and ACL fields on every chunkInternal tools; many tiny tenants

Principles:

  • Enforce at retrieval, server-side, from the authenticated identity - never from a parameter the client or the model can set.
  • Fail closed: if the user's permissions can't be resolved, return nothing.
  • Mirror source permissions (document ACLs, groups) into chunk metadata at ingest, and re-sync when permissions change - stale ACLs are a leak. How connectors, nested groups, ACL lag and late binding work is covered in Enterprise Data Integration.
  • Test it: include cross-tenant probes in the eval set and alert on any hit.
  • Cache keys must include the tenant and the permission set (see below).

Caching

LayerKey โ†’ valueHit rateRisks
Query embedding cachenormalised query text โ†’ vectorModerateNegligible
Retrieval cache(query, filters, index version) โ†’ candidate idsModerateStale after index updates - include index version in key
Exact answer cache(normalised query, tenant, permissions, index version, prompt version) โ†’ answerLow-moderate; high for FAQ-like trafficStaleness - invalidate on document change
Semantic answer cachenearest cached query above a similarity threshold โ†’ answerHigherFalse hits: "refund policy for Enterprise" and "for Starter" may be 0.95 similar but need different answers; cross-tenant leaks if the key omits the tenant
Prompt cache (provider)Stable prompt prefix โ†’ KV stateAutomaticSee Prompt Caching & Cost

Implement a semantic cache with a real vector index (not a scan over cache keys), partition it by tenant and permissions, use a conservative threshold tuned on labelled near-duplicate pairs, and keep TTLs short. For many products, an exact cache plus prompt caching delivers most of the savings with none of the false-hit risk.


Graceful Degradation

flowchart TD
    VF["๐Ÿ—„๏ธ Vector search down"] --> VF1["โ†ช๏ธ BM25-only retrieval"]
    RF["๐ŸŽฏ Reranker down or slow"] --> RF1["โ†ช๏ธ Serve first-stage order<br/>(tighter k)"]
    LF["๐Ÿค– LLM rate-limited / down"] --> LF1["โ†ช๏ธ Fallback model or provider"]
    LF1 --> LF2["โ†ช๏ธ Show retrieved passages<br/>without a generated answer"]
    EF["๐Ÿงฌ Embedding API down"] --> EF1["โ†ช๏ธ Cached embeddings or BM25"]

    style VF fill:#e8e0d4,stroke:#c8b89a
    style RF fill:#e8e0d4,stroke:#c8b89a
    style LF fill:#e8e0d4,stroke:#c8b89a
    style EF fill:#e8e0d4,stroke:#c8b89a

Wrap each dependency with timeouts and circuit breakers, and make the degraded mode explicit to users ("showing matching documents; answer generation is temporarily unavailable") rather than silently answering from the model's own knowledge.


Worked Design: Enterprise Knowledge Assistant

Prompt: 10M documents across HR, Legal and Engineering; 5,000 employees; answers with citations; p95 under 3 s; daily updates; department-level access control.

DecisionChoice and reason
Load5,000 users ร— 20 queries/day โ‰ˆ 100K/day โ‰ˆ 3.5 QPS average over 8 working hours, ~20 QPS peak
Chunks and index~12M chunks ร— 768 dims โ‰ˆ 37 GB float32 โ†’ int8 HNSW (~9 GB) with rescoring; hybrid BM25 in the same engine
Access controldepartment and document ACL groups on every chunk; filter from the SSO identity; fail closed
RetrievalHybrid top 50 โ†’ reranker top 6 โ†’ mid-size LLM with prompt caching of the instructions
FreshnessNightly batch plus event-driven updates for high-churn spaces; blue/green rebuild for model changes
Quality gates300-question golden set per department (incl. unanswerable and cross-department probes); retrieval metrics on every index change; groundedness sampled in production
DegradationBM25-only fallback; passages-only mode if the LLM is unavailable
Estimated latency~40 ms retrieval + ~150 ms rerank + ~800 ms TTFT โ†’ first token โ‰ˆ 1 s; full answer โ‰ˆ 2.5 s

Check Yourself

Check yourself
0 / 4 answered
  1. Where should the tenant/ACL filter be applied in a multi-tenant RAG system?
  2. A semantic answer cache with a 0.95 threshold returns Enterprise refund terms to a Starter-plan question. What is the root problem?
  3. In a typical RAG request, which stage usually dominates latency?
  4. How do you switch to a new embedding model without downtime?

Exercises

Exercise - Size and cost a design

Design RAG for 50M support articles and tickets (โ‰ˆ80M chunks, 1,024-dim embeddings), 200 QPS peak, p95 first token under 1.5 s, per-customer isolation for 3,000 customers. Specify index technology and memory, sharding and replication, the isolation model, the caching layers and the latency budget.

Solution

Vectors: 80M ร— 1,024 ร— 4 โ‰ˆ 328 GB float32 โ†’ binary in RAM (~10 GB) + int8 or float on SSD for rescoring, or a managed/distributed engine; shard by customer group so filtered queries hit few shards; replicas sized for 200 QPS with headroom. Isolation: namespaces per customer (or per group of small customers) plus mandatory ACL filters. Caching: embedding and exact-answer caches keyed by customer; provider prompt caching; semantic cache only for curated FAQ intents. Budget: ~50 ms retrieval, ~100 ms GPU rerank, ~800 ms TTFT with a mid-size model and short context.

Exercise - Failure drill

For the worked design above, write the runbook entries for (a) reranker latency doubling, (b) the vector index returning stale results after a failed nightly job, (c) a cross-department probe in the eval set returning a restricted chunk.

Solution

(a) Circuit breaker trips to first-stage order with smaller k; alert; investigate GPU capacity. (b) Index-freshness metric (newest indexed timestamp) alerts; rerun ingestion from the queue; if needed serve with a visible freshness warning. (c) Sev-1: disable the assistant for that department or all, audit the ACL sync for affected documents, fix, re-run the full cross-tenant probe suite before re-enabling.

Study Notes

Must-know:

  • Design flow: requirements โ†’ two pipelines โ†’ components โ†’ numbers โ†’ trade-offs โ†’ operations
  • Latency is dominated by the LLM; budget each stage; parallelise retrieval branches; stream
  • Index sizing: N ร— d ร— 4 bytes float32; compress, shard, replicate; blue/green rebuilds for model changes
  • Freshness: event-driven delete-then-insert with idempotent workers, or batch
  • Multi-tenancy: per-tenant index, namespaces, or shared + mandatory filters; enforce server-side from identity; fail closed; test with probes
  • Caches: embedding, retrieval, exact answer, semantic answer (false hits - key by tenant/permissions, conservative thresholds), provider prompt cache
  • Degrade explicitly: BM25 fallback, skip reranker, fallback model, passages-only mode

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท