RAG System Design
Designing a production RAG system means turning requirements - corpus size, query volume, latency, freshness, access control and quality targets - into concrete choices for ingestion, indexing, retrieval, generation, caching and failure handling, with the numbers to justify them.
- Run a structured RAG design discussion from requirements to trade-offs
- Build a per-stage latency budget and identify which stages to optimise
- Size an index and choose sharding, replication and freshness strategies for 10M-100M+ chunks
- Design multi-tenancy and access control that fail closed
- Specify caching layers and graceful degradation, including the risks of semantic caching
A Framework for the Design Discussion
| Step | Questions to settle |
|---|---|
| 1. Requirements | Corpus size and growth? Document types? QPS (average, peak)? Latency SLO (p95 TTFT and total)? Freshness (minutes, hours, daily)? Who may see what? Quality bar and cost per query? |
| 2. Two pipelines | Ingest (offline, event-driven or batch) and query (online) - draw both |
| 3. Components | Parsing and chunking; embedding model; index and store; retrieval (hybrid?), reranker; generator; citations |
| 4. Numbers | Index size, latency budget, cost per query, throughput per component |
| 5. Trade-offs | Latency vs quality (reranker, query rewriting), cost vs quality (model size, context size), freshness vs complexity |
| 6. Operations | Evaluation gates, monitoring, failure modes, security (RAG in Production) |
Reference Architecture
flowchart TD
subgraph ING["๐ฅ Ingestion (async)"]
SRC["๐๏ธ Sources<br/>object storage, wikis, tickets, DBs"] --> EV["๐ฌ Change events / scheduler"]
EV --> Q1["๐งพ Work queue"]
Q1 --> W["โ๏ธ Workers<br/>parse โ chunk โ enrich โ embed"]
W --> IDX[("๐๏ธ Hybrid index<br/>vectors + BM25 + metadata/ACL")]
W --> DOCS[("๐ฆ Document store<br/>full text, versions")]
end
subgraph QRY["๐ Query (sync)"]
U["๐ง User + identity"] --> API["๐ช API / gateway<br/>auth, rate limits"]
API --> QP["โ๏ธ Query processing<br/>rewrite, filters, route"]
QP --> RET["๐ก Hybrid retrieval<br/>ACL-filtered, top 50"]
RET --> RR["๐ฏ Reranker โ top 5-8"]
RR --> GEN["๐ค LLM, streaming<br/>grounded + cited"]
GEN --> OUT(["๐ฌ Answer + citations"])
end
IDX -.-> RET
DOCS -.-> RR
OUT -.-> FB["๐ Logs, feedback,<br/>sampled evals"]
style ING fill:#e8e2d9,stroke:#ccc4b8
style QRY fill:#d8dfe8,stroke:#b0bac8
style FB fill:#dde4dc,stroke:#b0c4b0
Latency Budget
A p95 target of about 3 seconds end to end, with streaming so the user sees text in about 1 second:
| Stage | Typical p95 | Levers |
|---|---|---|
| Auth, routing, filters | 10-30 ms | - |
| Query rewrite (LLM) | 200-800 ms | Skip for standalone queries; use a small model |
| Query embedding | 10-50 ms | Self-host or cache |
| Hybrid retrieval (ANN + BM25) | 10-80 ms | Index tuning, replicas, filter selectivity |
| Reranking 50 candidates | 30-300 ms | Smaller or GPU-hosted reranker; fewer candidates |
| LLM time to first token | 300-1,500 ms | Smaller model, shorter context, prompt caching, lower reasoning effort |
| LLM generation | 1-5 s | Stream it; concise answers |
Rules of thumb: the LLM dominates, so context size, model choice and caching matter more than shaving milliseconds from the vector search; run retrieval branches in parallel; and every optional LLM step (rewrite, grading, reranking with an LLM) needs to earn its latency on the eval.
Scaling the Index
Sizing. 100M documents ร 1.2 chunks ร 768 dimensions ร 4 bytes โ 369 GB of float32 vectors - more than a single machine's RAM for an in-memory HNSW index. Options, from Embeddings & Vector Search: int8 (โ92 GB) or binary with rescoring (โ12 GB in RAM), IVF-PQ or DiskANN, or a distributed managed service.
| Concern | Approach |
|---|---|
| Throughput | Replicate the index; queries are read-only and scale horizontally |
| Size | Shard; each shard searched in parallel, results merged (top-k per shard, then global top-k) |
| Shard key | Hash sharding spreads load evenly but every query fans out to all shards; sharding by tenant or domain lets filtered queries hit one shard |
| Freshness | Event-driven upserts: on change, delete the document's chunks by doc_id, then insert new ones; idempotent workers; a dead-letter queue for failures |
| Big rebuilds (new embedding model, new chunking) | Build a new index alongside, evaluate, switch traffic with an alias (blue/green), keep the old one for rollback |
| Ingestion throughput | Batch embedding calls, parallel workers behind a queue, back-pressure on API rate limits |
Multi-Tenancy and Access Control
Retrieval must never return a chunk the user isn't allowed to see - once a chunk is in the prompt, the model may repeat it.
| Isolation model | How | Use when |
|---|---|---|
| Index per tenant | Separate collections/indexes | Strict contractual or regulatory isolation; few, large tenants |
| Namespace / partition per tenant | Engine-level partitions in a shared cluster | Many tenants of varied size - the common SaaS default |
| Shared index + mandatory filter | tenant_id and ACL fields on every chunk | Internal tools; many tiny tenants |
Principles:
- Enforce at retrieval, server-side, from the authenticated identity - never from a parameter the client or the model can set.
- Fail closed: if the user's permissions can't be resolved, return nothing.
- Mirror source permissions (document ACLs, groups) into chunk metadata at ingest, and re-sync when permissions change - stale ACLs are a leak. How connectors, nested groups, ACL lag and late binding work is covered in Enterprise Data Integration.
- Test it: include cross-tenant probes in the eval set and alert on any hit.
- Cache keys must include the tenant and the permission set (see below).
Caching
| Layer | Key โ value | Hit rate | Risks |
|---|---|---|---|
| Query embedding cache | normalised query text โ vector | Moderate | Negligible |
| Retrieval cache | (query, filters, index version) โ candidate ids | Moderate | Stale after index updates - include index version in key |
| Exact answer cache | (normalised query, tenant, permissions, index version, prompt version) โ answer | Low-moderate; high for FAQ-like traffic | Staleness - invalidate on document change |
| Semantic answer cache | nearest cached query above a similarity threshold โ answer | Higher | False hits: "refund policy for Enterprise" and "for Starter" may be 0.95 similar but need different answers; cross-tenant leaks if the key omits the tenant |
| Prompt cache (provider) | Stable prompt prefix โ KV state | Automatic | See Prompt Caching & Cost |
Implement a semantic cache with a real vector index (not a scan over cache keys), partition it by tenant and permissions, use a conservative threshold tuned on labelled near-duplicate pairs, and keep TTLs short. For many products, an exact cache plus prompt caching delivers most of the savings with none of the false-hit risk.
Graceful Degradation
flowchart TD
VF["๐๏ธ Vector search down"] --> VF1["โช๏ธ BM25-only retrieval"]
RF["๐ฏ Reranker down or slow"] --> RF1["โช๏ธ Serve first-stage order<br/>(tighter k)"]
LF["๐ค LLM rate-limited / down"] --> LF1["โช๏ธ Fallback model or provider"]
LF1 --> LF2["โช๏ธ Show retrieved passages<br/>without a generated answer"]
EF["๐งฌ Embedding API down"] --> EF1["โช๏ธ Cached embeddings or BM25"]
style VF fill:#e8e0d4,stroke:#c8b89a
style RF fill:#e8e0d4,stroke:#c8b89a
style LF fill:#e8e0d4,stroke:#c8b89a
style EF fill:#e8e0d4,stroke:#c8b89a
Wrap each dependency with timeouts and circuit breakers, and make the degraded mode explicit to users ("showing matching documents; answer generation is temporarily unavailable") rather than silently answering from the model's own knowledge.
Worked Design: Enterprise Knowledge Assistant
Prompt: 10M documents across HR, Legal and Engineering; 5,000 employees; answers with citations; p95 under 3 s; daily updates; department-level access control.
| Decision | Choice and reason |
|---|---|
| Load | 5,000 users ร 20 queries/day โ 100K/day โ 3.5 QPS average over 8 working hours, ~20 QPS peak |
| Chunks and index | ~12M chunks ร 768 dims โ 37 GB float32 โ int8 HNSW (~9 GB) with rescoring; hybrid BM25 in the same engine |
| Access control | department and document ACL groups on every chunk; filter from the SSO identity; fail closed |
| Retrieval | Hybrid top 50 โ reranker top 6 โ mid-size LLM with prompt caching of the instructions |
| Freshness | Nightly batch plus event-driven updates for high-churn spaces; blue/green rebuild for model changes |
| Quality gates | 300-question golden set per department (incl. unanswerable and cross-department probes); retrieval metrics on every index change; groundedness sampled in production |
| Degradation | BM25-only fallback; passages-only mode if the LLM is unavailable |
| Estimated latency | ~40 ms retrieval + ~150 ms rerank + ~800 ms TTFT โ first token โ 1 s; full answer โ 2.5 s |
Check Yourself
- Where should the tenant/ACL filter be applied in a multi-tenant RAG system?
- A semantic answer cache with a 0.95 threshold returns Enterprise refund terms to a Starter-plan question. What is the root problem?
- In a typical RAG request, which stage usually dominates latency?
- How do you switch to a new embedding model without downtime?
Exercises
Design RAG for 50M support articles and tickets (โ80M chunks, 1,024-dim embeddings), 200 QPS peak, p95 first token under 1.5 s, per-customer isolation for 3,000 customers. Specify index technology and memory, sharding and replication, the isolation model, the caching layers and the latency budget.
Solution
Vectors: 80M ร 1,024 ร 4 โ 328 GB float32 โ binary in RAM (~10 GB) + int8 or float on SSD for rescoring, or a managed/distributed engine; shard by customer group so filtered queries hit few shards; replicas sized for 200 QPS with headroom. Isolation: namespaces per customer (or per group of small customers) plus mandatory ACL filters. Caching: embedding and exact-answer caches keyed by customer; provider prompt caching; semantic cache only for curated FAQ intents. Budget: ~50 ms retrieval, ~100 ms GPU rerank, ~800 ms TTFT with a mid-size model and short context.
For the worked design above, write the runbook entries for (a) reranker latency doubling, (b) the vector index returning stale results after a failed nightly job, (c) a cross-department probe in the eval set returning a restricted chunk.
Solution
(a) Circuit breaker trips to first-stage order with smaller k; alert; investigate GPU capacity. (b) Index-freshness metric (newest indexed timestamp) alerts; rerun ingestion from the queue; if needed serve with a visible freshness warning. (c) Sev-1: disable the assistant for that department or all, audit the ACL sync for affected documents, fix, re-run the full cross-tenant probe suite before re-enabling.
Study Notes
Must-know:
- Design flow: requirements โ two pipelines โ components โ numbers โ trade-offs โ operations
- Latency is dominated by the LLM; budget each stage; parallelise retrieval branches; stream
- Index sizing: N ร d ร 4 bytes float32; compress, shard, replicate; blue/green rebuilds for model changes
- Freshness: event-driven delete-then-insert with idempotent workers, or batch
- Multi-tenancy: per-tenant index, namespaces, or shared + mandatory filters; enforce server-side from identity; fail closed; test with probes
- Caches: embedding, retrieval, exact answer, semantic answer (false hits - key by tenant/permissions, conservative thresholds), provider prompt cache
- Degrade explicitly: BM25 fallback, skip reranker, fallback model, passages-only mode
References
- Gao et al., Retrieval-Augmented Generation for Large Language Models: A Survey (2023)
- Barnett et al., Seven Failure Points When Engineering a Retrieval Augmented Generation System (CAIN 2024)
- Google Cloud, RAG infrastructure for generative AI using Agent Platform and Vector Search
- AWS, Amazon Bedrock Knowledge Bases
Last reviewed: 2026-09