Contents
Map

Quiz · 12 · RAG

56 questions from 14 pages

These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →

RAG Fundamentals

Check yourself
0 / 4 answered
  1. After upgrading the embedding model used for queries, every answer becomes irrelevant but no errors appear. What happened?
  2. Which question is RAG over documents least suited to?
  3. A 40-page product manual changes monthly and every answer must cite the page. Which approach is simplest?
  4. Why does fine-tuning not replace RAG for fast-changing facts?
Check yourself
0 / 4 answered
  1. How many bytes per vector does IVF-PQ with 96 sub-quantizers of 8 bits use for 768-dimensional embeddings, and what compression is that versus float32?
  2. Your model card says to prefix queries with 'query: ' and passages with 'passage: '. You forget the prefixes. What happens?
  3. Why is L2 distance equivalent to cosine similarity for ranking unit-length vectors?
  4. A model ranks first on the MTEB retrieval leaderboard. Why might it still lose to a lower-ranked model on your data?

Document Processing & Chunking

Check yourself
0 / 4 answered
  1. A chunk reads 'Revenue grew 3% versus the prior quarter.' Queries about 'ACME Q2 revenue' never retrieve it. Which technique targets exactly this problem?
  2. What does parent-child retrieval decouple?
  3. What is the best evidence-based default when choosing between semantic and recursive chunking?
  4. Why prepend the section heading path to each chunk?

Retrieval & Reranking

Check yourself
0 / 4 answered
  1. A document is ranked 1st by BM25 and not retrieved by dense search; another is ranked 4th by both. With RRF (k=60), which scores higher?
  2. Why are cross-encoders used for reranking rather than first-stage retrieval?
  3. Adding an off-the-shelf reranker to your pipeline slightly lowers nDCG@10 on your domain. What is the most likely explanation?
  4. Which query does BM25 handle better than a dense retriever, and why?

Grounded Generation & Citations

Check yourself
0 / 4 answered
  1. The retrieved passage says a drug's dose is 1200 mg (a typo; the correct dose is 120 mg), and the answer repeats 1200 mg with a citation. How is this failure classified?
  2. What does citation precision measure?
  3. Why can sending 10 passages instead of 3 lower answer accuracy even when the right passage is among them?
  4. How do you evaluate abstention properly?

RAG Evaluation

Check yourself
0 / 4 answered
  1. A query has one relevant document, ranked 2nd. What are MRR@10 and nDCG@10?
  2. Context recall is 0.92 but faithfulness is 0.55. Where is the problem?
  3. Why are LLM-generated synthetic test questions often too easy for retrieval?
  4. Can an answer be perfectly grounded and still wrong? Give an example.

Advanced RAG Patterns

Check yourself
0 / 4 answered
  1. What do Self-RAG's reflection tokens require that CRAG does not?
  2. What does Speculative RAG do?
  3. Which question most clearly calls for GraphRAG global search rather than vector RAG?
  4. Why does ColPali not need OCR?

Agentic & Deep-Research RAG

Check yourself
0 / 4 answered
  1. Which question most justifies agentic RAG over a pipeline?
  2. According to Anthropic's analysis of its research system, what explained most of the variance in BrowseComp performance?
  3. Your agent sometimes performs 40 searches on unanswerable questions. What are the two most direct fixes?
  4. Why do deep-research systems use sub-agents with separate contexts rather than one long-running agent?

Long Context vs RAG vs CAG

Check yourself
0 / 4 answered
  1. With hosted APIs, what is cache-augmented generation in practice?
  2. What did Li et al.'s Self-Route do?
  3. Why is CAG a poor fit for a multi-tenant assistant where each customer sees only their own documents?
  4. A 64-layer model with 8 KV heads of dimension 128 caches in bf16. Roughly how much KV memory does a 100K-token cached corpus need?

Enterprise Data Integration

Check yourself
0 / 5 answered
  1. Why should webhooks or change notifications be combined with periodic delta syncs and reconciliation?
  2. Where should the user's permission filter be applied in a permission-aware RAG system?
  3. A document's permissions were tightened 10 minutes ago, but the assistant still quotes it to a user who lost access. What happened, and how do you fix it?
  4. What is the oversharing problem in enterprise assistants?
  5. A user exercises their right to erasure. Which copies of their data must the RAG system remove?

RAG System Design

Check yourself
0 / 4 answered
  1. Where should the tenant/ACL filter be applied in a multi-tenant RAG system?
  2. A semantic answer cache with a 0.95 threshold returns Enterprise refund terms to a Starter-plan question. What is the root problem?
  3. In a typical RAG request, which stage usually dominates latency?
  4. How do you switch to a new embedding model without downtime?

RAG in Production

Check yourself
0 / 4 answered
  1. Which is the best autoscaling signal for a self-hosted LLM used by the RAG service?
  2. Why should embeddings be protected like the source text?
  3. A new wiki page is retrieved for hundreds of unrelated queries the day after it was created and answers shift toward recommending one vendor. What is the most likely attack?
  4. What must be logged with each answer to reproduce it later?

Managed RAG on Cloud Platforms

Check yourself
0 / 4 answered
  1. Your team's Gemini RAG code imports vertexai.generative_models and now fails. What changed?
  2. Which requirement most strongly argues for a custom build over a fully managed knowledge base?
  3. In Bedrock's retrieve_and_generate, where do you control hybrid vs semantic search?
  4. Why insist on access to retrieved chunks and scores from a managed RAG service?

Retrieval Evaluation

Check yourself
0 / 3 answered
  1. Hybrid has Recall@100 = 0.955 but nDCG@10 = 0.690. What does that tell you about where to invest next?
  2. Why does the reranker in this lab slightly lower nDCG@10?
  3. Why is matching published BEIR numbers for BM25 a useful check?