Quiz · 01 · LLM Foundations
38 questions from 8 pages
These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
LLM Fundamentals
Check yourself
0 / 5 answered
- With OpenAI's o200k_base tokenizer, how is "9.11" split?
- A 128K-context model receives a 100K-token prompt. What is the most it can generate?
- What does FlashAttention change about long-context cost?
- Why might two identical requests at temperature 0 return different text?
- When does top-p (nucleus) sampling keep only a few candidate tokens?
Tokenization
Check yourself
0 / 5 answered
- Why does byte-level BPE never need an 'unknown' token?
- Moving from a 32K to a 256K vocabulary, what typically happens?
- In the table above, the Hindi sentence takes 20 tokens with cl100k_base and 9 with o200k_base. What practical effects does that have for a Hindi-language product?
- A model fails to say whether 9.11 or 9.9 is larger. What is a tokenization-based explanation?
- Why must user-provided text be encoded without special-token parsing?
Transformer Architecture
Check yourself
0 / 5 answered
- Why do Pre-LN transformers train more stably than Post-LN ones?
- A model's config has d_model=4096, 32 query heads, 8 KV heads, d_head=128. How many parameters are in one layer's K projection?
- Why can't a model with learned absolute position embeddings handle sequences longer than max_seq_len?
- What does RMSNorm drop compared with LayerNorm?
- Where do most of a dense transformer's parameters live?
Attention Mechanisms
Check yourself
0 / 5 answered
- Why divide Q·Kᵀ by √d_k?
- In a causal decoder, which positions can token i attend to?
- A model has 32 query heads and 8 KV heads (GQA). How much smaller is its KV cache than with full MHA?
- What does FlashAttention change?
- Why can causal masking train on every position of a sequence in one forward pass?
Model Architecture Types
Check yourself
0 / 5 answered
- Why can't BERT generate free text?
- Mixtral 8x7B has about 47B parameters, not 56B. Why?
- What does a mixture-of-experts model save compared with a dense model of the same total size?
- What distinguishes a reasoning model from an ordinary chat model with the same architecture?
- You need fast, cheap intent classification for 50 million support messages a day. Which architecture class do you start with, and why?
Modern Architectures
Check yourself
0 / 4 answered
- A 671B-parameter MoE model activates 37B parameters per token. Compared with a 37B dense model, what is roughly true?
- What does MLA cache for each token, instead of per-head keys and values?
- Why do models like Gemma 3 interleave sliding-window layers with a few global-attention layers rather than using only sliding windows?
- Why are most production state-space models hybrids that keep some attention layers?
Failure Modes & Tricky Issues
Check yourself
0 / 6 answered
- Why does LoRA reduce catastrophic forgetting without eliminating it?
- Retrieved context contains the answer, but it sits in the middle of 20 documents and the model misses it. Which fix addresses the mechanism?
- A user pushes back on a correct answer and the model changes its answer to agree. What is this, and where does it come from?
- Which is the reliable fix for arithmetic and exact string operations?
- Why is the intrinsic vs extrinsic distinction useful when evaluating a RAG system?
- RAG reduces hallucination - why doesn't it eliminate it?
Model Landscape
Check yourself
0 / 3 answered
- Two models score within a point of each other on your evaluation. Model A costs half as much per token but needs, on average, 2.5 attempts to finish a task; Model B finishes in one. Which is cheaper per completed task?
- You must self-host on 8× 80 GB GPUs. Which fact about an MoE model decides whether it fits?
- Why does this course keep current model names on one page instead of in every note?