Contents
Map

Quiz · 01 · LLM Foundations

38 questions from 8 pages

These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →

LLM Fundamentals

Check yourself
0 / 5 answered
  1. With OpenAI's o200k_base tokenizer, how is "9.11" split?
  2. A 128K-context model receives a 100K-token prompt. What is the most it can generate?
  3. What does FlashAttention change about long-context cost?
  4. Why might two identical requests at temperature 0 return different text?
  5. When does top-p (nucleus) sampling keep only a few candidate tokens?

Tokenization

Check yourself
0 / 5 answered
  1. Why does byte-level BPE never need an 'unknown' token?
  2. Moving from a 32K to a 256K vocabulary, what typically happens?
  3. In the table above, the Hindi sentence takes 20 tokens with cl100k_base and 9 with o200k_base. What practical effects does that have for a Hindi-language product?
  4. A model fails to say whether 9.11 or 9.9 is larger. What is a tokenization-based explanation?
  5. Why must user-provided text be encoded without special-token parsing?

Transformer Architecture

Check yourself
0 / 5 answered
  1. Why do Pre-LN transformers train more stably than Post-LN ones?
  2. A model's config has d_model=4096, 32 query heads, 8 KV heads, d_head=128. How many parameters are in one layer's K projection?
  3. Why can't a model with learned absolute position embeddings handle sequences longer than max_seq_len?
  4. What does RMSNorm drop compared with LayerNorm?
  5. Where do most of a dense transformer's parameters live?

Attention Mechanisms

Check yourself
0 / 5 answered
  1. Why divide Q·Kᵀ by √d_k?
  2. In a causal decoder, which positions can token i attend to?
  3. A model has 32 query heads and 8 KV heads (GQA). How much smaller is its KV cache than with full MHA?
  4. What does FlashAttention change?
  5. Why can causal masking train on every position of a sequence in one forward pass?

Model Architecture Types

Check yourself
0 / 5 answered
  1. Why can't BERT generate free text?
  2. Mixtral 8x7B has about 47B parameters, not 56B. Why?
  3. What does a mixture-of-experts model save compared with a dense model of the same total size?
  4. What distinguishes a reasoning model from an ordinary chat model with the same architecture?
  5. You need fast, cheap intent classification for 50 million support messages a day. Which architecture class do you start with, and why?

Modern Architectures

Check yourself
0 / 4 answered
  1. A 671B-parameter MoE model activates 37B parameters per token. Compared with a 37B dense model, what is roughly true?
  2. What does MLA cache for each token, instead of per-head keys and values?
  3. Why do models like Gemma 3 interleave sliding-window layers with a few global-attention layers rather than using only sliding windows?
  4. Why are most production state-space models hybrids that keep some attention layers?

Failure Modes & Tricky Issues

Check yourself
0 / 6 answered
  1. Why does LoRA reduce catastrophic forgetting without eliminating it?
  2. Retrieved context contains the answer, but it sits in the middle of 20 documents and the model misses it. Which fix addresses the mechanism?
  3. A user pushes back on a correct answer and the model changes its answer to agree. What is this, and where does it come from?
  4. Which is the reliable fix for arithmetic and exact string operations?
  5. Why is the intrinsic vs extrinsic distinction useful when evaluating a RAG system?
  6. RAG reduces hallucination - why doesn't it eliminate it?

Model Landscape

Check yourself
0 / 3 answered
  1. Two models score within a point of each other on your evaluation. Model A costs half as much per token but needs, on average, 2.5 attempts to finish a task; Model B finishes in one. Which is cheaper per completed task?
  2. You must self-host on 8× 80 GB GPUs. Which fact about an MoE model decides whether it fits?
  3. Why does this course keep current model names on one page instead of in every note?