These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
Check yourself 0 / 4 answered
After upgrading the embedding model used for queries, every answer becomes irrelevant but no errors appear. What happened?
A The corpus was embedded with the old model, so query and document vectors are in different spaces B The LLM changed C The chunk size was too small D The reranker is misconfigured Which question is RAG over documents least suited to?
A What was total revenue by region last quarter? B What does our refund policy say about annual plans? C Which runbook covers database failover? D What did the 2025 security audit recommend? A 40-page product manual changes monthly and every answer must cite the page. Which approach is simplest?
A Put the whole manual in a cached prompt and ask for quoted citations B Fine-tune a model on the manual each month C Build a GraphRAG index D Train a custom embedding model Why does fine-tuning not replace RAG for fast-changing facts?
Show answer
Check yourself 0 / 4 answered
How many bytes per vector does IVF-PQ with 96 sub-quantizers of 8 bits use for 768-dimensional embeddings, and what compression is that versus float32?
A 96 bytes, 32× smaller than 3,072 bytes B 768 bytes, 4× smaller C 8 bytes, 384× smaller D 12 bytes, 256× smaller Your model card says to prefix queries with 'query: ' and passages with 'passage: '. You forget the prefixes. What happens?
A Retrieval still works but is measurably worse, because the model was trained to encode the two roles differently B The API returns an error C Nothing - prefixes are cosmetic D Embedding dimension changes Why is L2 distance equivalent to cosine similarity for ranking unit-length vectors?
A Because ‖q − d‖² = 2 − 2·cos(q, d) when ‖q‖ = ‖d‖ = 1 B Because both ignore direction C Because L2 is always larger D They are not equivalent A model ranks first on the MTEB retrieval leaderboard. Why might it still lose to a lower-ranked model on your data?
Show answer
Check yourself 0 / 4 answered
A chunk reads 'Revenue grew 3% versus the prior quarter.' Queries about 'ACME Q2 revenue' never retrieve it. Which technique targets exactly this problem?
A Contextual retrieval: prepend an LLM-written context naming the company and period before embedding and BM25 indexing B Increasing chunk overlap C Switching from cosine to dot product D Lowering the temperature What does parent-child retrieval decouple?
A The unit you search over (small chunks) from the unit you give the model (larger parents) B Dense from sparse retrieval C The embedding model from the LLM D Metadata from text What is the best evidence-based default when choosing between semantic and recursive chunking?
A Start with structure-aware recursive chunking and tune size on an eval; semantic chunking has not shown consistent gains B Always use semantic chunking - it is state of the art C Always use 1,000-character fixed chunks D Never overlap chunks Why prepend the section heading path to each chunk?
Show answer
Check yourself 0 / 4 answered
A document is ranked 1st by BM25 and not retrieved by dense search; another is ranked 4th by both. With RRF (k=60), which scores higher?
A The one ranked 4th by both (2/64 ≈ 0.031 vs 1/61 ≈ 0.016) B The one ranked 1st by BM25 C They tie D It depends on the BM25 score Why are cross-encoders used for reranking rather than first-stage retrieval?
A They score each (query, passage) pair jointly, so nothing can be precomputed - too slow to run over a whole corpus B They have lower accuracy than bi-encoders C They can't handle long passages D They only support English Adding an off-the-shelf reranker to your pipeline slightly lowers nDCG@10 on your domain. What is the most likely explanation?
A The reranker's training data (e.g. web search) doesn't match your domain B Rerankers always hurt precision C RRF is incompatible with rerankers D The candidate depth was too large Which query does BM25 handle better than a dense retriever, and why?
Show answer
Check yourself 0 / 4 answered
The retrieved passage says a drug's dose is 1200 mg (a typo; the correct dose is 120 mg), and the answer repeats 1200 mg with a citation. How is this failure classified?
A Faithful to a bad source - a source-quality problem, not an unfaithful generation B Context ignored C Embellishment D A retrieval failure What does citation precision measure?
A Whether each cited passage actually supports the statement it's attached to B Whether every statement has a citation C How many passages were retrieved D Whether citations are formatted correctly Why can sending 10 passages instead of 3 lower answer accuracy even when the right passage is among them?
A Irrelevant passages distract the model and push evidence toward the middle of the context B The model can only read 3 documents C Citations stop working above 5 documents D The prompt cache is invalidated How do you evaluate abstention properly?
Show answer
Check yourself 0 / 4 answered
A query has one relevant document, ranked 2nd. What are MRR@10 and nDCG@10?
A MRR = 0.5, nDCG = 1/log2(3) ≈ 0.63 B MRR = 0.5, nDCG = 0.5 C MRR = 1, nDCG = 1 D MRR = 2, nDCG = 0.63 Context recall is 0.92 but faithfulness is 0.55. Where is the problem?
A Generation - the evidence is retrieved but the answer isn't sticking to it B Retrieval recall C Chunking D The embedding model Why are LLM-generated synthetic test questions often too easy for retrieval?
A They reuse the source chunk's wording, so lexical and semantic overlap with the target chunk is unrealistically high B They are too long C They are always multi-hop D They have no reference answers Can an answer be perfectly grounded and still wrong? Give an example.
Show answer
Check yourself 0 / 4 answered
What do Self-RAG's reflection tokens require that CRAG does not?
A A generator fine-tuned to emit them B A web search API C A knowledge graph D Token probabilities What does Speculative RAG do?
A A small model drafts several answers from different document subsets in parallel; a large model verifies and selects one B It pre-fetches documents the user might ask about next C It speculatively decodes tokens with a draft model D It caches retrieval results Which question most clearly calls for GraphRAG global search rather than vector RAG?
A What are the recurring root causes across 3,000 incident post-mortems? B What is the SLA for the Enterprise plan? C How do I reset my password? D What is error code E4013? Why does ColPali not need OCR?
Show answer
Check yourself 0 / 4 answered
Which question most justifies agentic RAG over a pipeline?
A Which of our vendors were affected by regulations that took effect after our last audit? B What is the refund window for Enterprise plans? C Where is the VPN setup guide? D What does error E4013 mean? According to Anthropic's analysis of its research system, what explained most of the variance in BrowseComp performance?
A Token usage B The embedding model C Number of tools available D Temperature Your agent sometimes performs 40 searches on unanswerable questions. What are the two most direct fixes?
A An explicit search budget told to the model, and an instruction (plus eval items) for stopping and reporting what's missing B A larger context window and higher temperature C Removing the search tool D Switching to BM25 Why do deep-research systems use sub-agents with separate contexts rather than one long-running agent?
Show answer
Check yourself 0 / 4 answered
With hosted APIs, what is cache-augmented generation in practice?
A Prompt caching of a prefix containing the whole corpus, reused across questions B Caching previous answers to identical questions C Caching embeddings of the corpus D Fine-tuning the model on the corpus What did Li et al.'s Self-Route do?
A Tried RAG first and fell back to long context only when the model judged the retrieved chunks insufficient B Routed queries between two embedding models C Used long context for all queries D Trained a retriever end to end Why is CAG a poor fit for a multi-tenant assistant where each customer sees only their own documents?
A The cached prefix would differ per tenant, so it can't be shared, and per-user access filtering is what RAG provides B CAG can't handle PDFs C CAG requires fine-tuning D CAG only works with small models A 64-layer model with 8 KV heads of dimension 128 caches in bf16. Roughly how much KV memory does a 100K-token cached corpus need?
Show answer
Check yourself 0 / 5 answered
Why should webhooks or change notifications be combined with periodic delta syncs and reconciliation?
A Webhooks are slower than polling B Notifications can be dropped or delayed, so a scheduled sync and a full comparison catch missed changes and deletes C Delta APIs are required by law D Webhooks can't carry document content Where should the user's permission filter be applied in a permission-aware RAG system?
A In the prompt: tell the model not to reveal restricted content B After retrieval, by removing unauthorized chunks from the top-k C Inside the vector and keyword queries as a metadata filter on ACL principals, failing closed if identity can't be resolved D Only in the user interface A document's permissions were tightened 10 minutes ago, but the assistant still quotes it to a user who lost access. What happened, and how do you fix it?
Show answer What is the oversharing problem in enterprise assistants?
Show answer A user exercises their right to erasure. Which copies of their data must the RAG system remove?
Show answer
Check yourself 0 / 4 answered
Where should the tenant/ACL filter be applied in a multi-tenant RAG system?
A In the retrieval query, server-side, derived from the authenticated identity B In the system prompt, telling the model not to reveal other tenants' data C After generation, by redacting the answer D In the client application A semantic answer cache with a 0.95 threshold returns Enterprise refund terms to a Starter-plan question. What is the root problem?
A Embedding similarity doesn't capture the one attribute that changes the answer; the cache threshold/key design is unsafe for this query class B The TTL is too long C The embedding model is too small D The LLM hallucinated In a typical RAG request, which stage usually dominates latency?
A LLM prefill and generation B ANN search C BM25 D Auth How do you switch to a new embedding model without downtime?
Show answer
Check yourself 0 / 4 answered
Which is the best autoscaling signal for a self-hosted LLM used by the RAG service?
A Request queue depth / KV-cache usage B GPU utilisation C CPU utilisation of the API pods D Number of documents in the index Why should embeddings be protected like the source text?
A Inversion attacks can reconstruct much of the original text from its embedding B Embeddings are larger than the text C Embeddings contain the model weights D They can't be deleted A new wiki page is retrieved for hundreds of unrelated queries the day after it was created and answers shift toward recommending one vendor. What is the most likely attack?
A Corpus poisoning - a document crafted to be retrieved and to steer answers B Embedding inversion C A denial-of-service attack D Model drift What must be logged with each answer to reproduce it later?
Show answer
Check yourself 0 / 4 answered
Your team's Gemini RAG code imports vertexai.generative_models and now fails. What changed?
A Google removed the vertexai.generative_models module in June 2026; model calls moved to the google-genai SDK B Gemini no longer supports RAG C RAG Engine was discontinued D The code needs a newer Python version Which requirement most strongly argues for a custom build over a fully managed knowledge base?
A Retrieval quality is a product differentiator that needs custom ranking and domain models B Documents live in SharePoint C Users are internal employees D You need answers with citations In Bedrock's retrieve_and_generate, where do you control hybrid vs semantic search?
A vectorSearchConfiguration.overrideSearchType (HYBRID or SEMANTIC) B The modelArn C The knowledge base name D The input text Why insist on access to retrieved chunks and scores from a managed RAG service?
Show answer
Check yourself 0 / 3 answered
Hybrid has Recall@100 = 0.955 but nDCG@10 = 0.690. What does that tell you about where to invest next?
A The right documents are almost always in the top 100; improving the ordering (a better, in-domain reranker) is the lever B First-stage recall is the bottleneck C The metrics are inconsistent D Chunking must be changed Why does the reranker in this lab slightly lower nDCG@10?
A It was trained on web search (MS MARCO), which differs from scientific claim verification B Rerankers always lower nDCG C RRF output can't be reranked D 50 candidates is too many Why is matching published BEIR numbers for BM25 a useful check?
Show answer