Contents
Map

01 · LLM Foundations

Q&A Review Bank

View as:

LLM Models - Q&A Review Bank

47 curated Q&A pairs. Tags: [Easy] = conceptual recall, [Medium] = design decisions / tradeoffs, [Hard] = math-level, system design, or tricky edge cases.

Learning objectives 65 min
By the end of this page you will be able to:
  • Answer each question from memory before revealing the answer, across: Fundamentals and Tokens; Transformer Architecture; Attention Mechanisms; Model Architecture Types; Failure Modes; Long Inputs; Architecture Design Decisions; Model Landscape Basics; Tokenization
  • Explain the reasoning behind each answer - the mechanism or trade-off - not only the fact
  • Identify the chapters you are weakest on and revisit them before the module quiz
Prerequisites
  • The concept notes of this module

Section 1: Fundamentals and Tokens

Q1 [Easy] What is a Large Language Model at its most fundamental level?

A next-token predictor. An LLM is a neural network trained to predict the most likely next token given all preceding tokens (Causal Language Modeling). All capabilities - reasoning, coding, summarization - emerge from this single objective applied at scale.


Q2 [Easy] What is a token and why do LLMs not operate on words?

A token is a subword unit produced by BPE/SentencePiece. Operating on words would require a vocabulary of hundreds of thousands of entries (all words across all languages). BPE produces a compact 32K–200K vocabulary that covers all text efficiently, represents rare words as subword pieces, and handles code/non-Latin scripts uniformly.


Q3 [Easy] How many tokens is roughly equivalent to one word in English?

1 token ≈ 0.75 words ≈ 4 characters. Roughly 100 tokens ≈ 75 words ≈ one short paragraph.


Q4 [Medium] Why does the model fail at "Is 9.11 greater than 9.9?"

Tokenization. "9.11" tokenizes as ["9", ".", "1", "1"] and "9.9" as ["9", ".", "9"]. The model sees digit token sequences, not numbers, and has no native integer comparison - it must reason about ordering from statistical patterns in training data. This is why LLMs should use Python tools for any numerical computation.


Q5 [Medium] What is the difference between temperature=0 and greedy decoding?

Mathematically equivalent. As temperature → 0, softmax concentrates all probability mass on the argmax token, producing the same result as greedy argmax. In practice, most APIs implement temperature=0 directly as argmax to avoid floating-point instability (dividing logits by ~0). The behavioral output is identical.


Q6 [Medium] Explain top-p (nucleus) sampling and why it's preferred over top-k.

Top-p considers the smallest set of tokens whose cumulative probability ≥ p, then samples from that nucleus. It's adaptive: when the model is confident (peaked distribution), the nucleus is small (few tokens considered); when uncertain, the nucleus is large. Top-k uses a fixed cutoff (always k tokens) regardless of distribution shape - it may include too many options when the model is certain, or too few when it's uncertain. Top-p adapts to the model's confidence level.


Q7 [Hard] Why does greedy decoding sometimes produce repetition loops, and how do you fix it?

Greedy decoding creates attractor states: once a repeated token sequence is in the context, the conditional probability of continuing the loop is higher than breaking out (the model has seen lots of repetitive text in training). repetition_penalty > 1.0 divides the logits of already-seen tokens, breaking the positive feedback loop. no_repeat_ngram_size=n explicitly forbids any n-gram from appearing twice. Sampling with temperature > 0 also naturally avoids deterministic attractors.


Section 2: Transformer Architecture

Q8 [Easy] What problem did transformers solve that RNNs couldn't?

Parallelization and long-range dependencies. RNNs process tokens sequentially (token 1 → token 2 → ...) so training cannot be parallelized. They also suffer from gradient vanishing for long-range dependencies. Transformers compute all pairwise token relationships in one parallel pass, enabling massive GPU parallelism and direct access to any position regardless of distance.


Q9 [Easy] What is a residual connection and why does it matter?

A residual connection adds the input directly to the output of a sublayer: output = x + Sublayer(x). It creates a "gradient highway" - gradients can flow directly from the output to any earlier layer without passing through all transformations. Without residuals, gradients vanish in deep networks (32+ layers). With residuals, deep transformers train stably.


Q10 [Medium] What is the difference between Pre-LN and Post-LN, and which is better?

Post-LN (original Transformer): output = LayerNorm(x + Sublayer(x)) - LayerNorm applied after the residual. Pre-LN (modern LLMs): output = x + Sublayer(LayerNorm(x)) - LayerNorm applied before the sublayer. Pre-LN is better: in Pre-LN, the residual path bypasses LayerNorm entirely, giving gradients a clean path to early layers. This makes training more stable without requiring careful learning rate warmup, enabling easier scaling to very deep networks.


Q11 [Medium] What is SwiGLU and why does LLaMA use it instead of ReLU in the FFN?

SwiGLU: FFN(x) = W2 · (SiLU(W1·x) ⊗ W3·x). It uses three weight matrices (vs two for ReLU-FFN) and a gating mechanism. SwiGLU consistently outperforms ReLU and GELU on language modeling benchmarks - the gating provides smoother activation and better gradient flow. LLaMA, Gemma, and Mistral all use SwiGLU.


Q12 [Hard] What is RoPE and how does it differ from learned absolute positional embeddings?

Learned absolute embeddings: a trainable lookup table [max_seq, d_model] - different embedding per absolute position. Cannot extrapolate beyond max_seq. RoPE (Rotary Position Embeddings): rotates Q and K vectors by an angle proportional to their absolute position. The dot product Q_i · K_j then encodes the relative position i-j naturally. Benefits: (1) relative position is what matters for language understanding, (2) RoPE can extrapolate to longer sequences than seen during training (via frequency scaling, e.g., YaRN), (3) no extra parameters.


Q13 [Medium] What does the FFN sublayer do that attention doesn't?

Attention routes information between tokens - it computes a weighted mixture of other tokens' value vectors. FFN applies the same two-layer MLP to each token independently (no cross-token interaction). FFN is thought to store factual associations - the "knowledge" of the model - while attention handles reasoning and information routing. The two sublayers have complementary roles.


Section 3: Attention Mechanisms

Q14 [Easy] Write the attention formula.

Attention(Q, K, V) = softmax(Q·Kᵀ / √d_k) · V


Q15 [Easy] Why do we scale Q·Kᵀ by √d_k?

The dot product of two d_k-dimensional random vectors has variance proportional to d_k. Without scaling, large d_k produces very large dot products → softmax becomes extremely peaked (near one-hot) → near-zero gradients → training stalls. Dividing by √d_k normalizes the variance to ~1, keeping softmax in its informative gradient range.


Q16 [Medium] Why does multi-head attention use separate W_Q/W_K/W_V per head rather than just splitting the embedding?

Separate projection matrices allow each head to compute a different linear transformation of the full d_model embedding - a genuinely different "view" of the input. Simply slicing d_model/h elements per head gives each head a different portion of the same embedding, which was computed as a unit by the preceding FFN and LayerNorm. The heads would see correlated, non-independent views. Separate projections allow truly independent perspectives: one head can focus on subject-verb agreement while another tracks coreference.


Q17 [Medium] What is the causal mask and why is it needed for training?

The causal mask sets all positions where j > i to -inf before softmax, preventing token i from attending to future tokens j > i. This is necessary because the model is trained to predict the next token - it must not see the answer. The mask enables parallel training: all positions can be predicted simultaneously in one forward pass, each using only its valid past context, instead of running N sequential forward passes for N positions.


Q18 [Hard] What is Flash Attention and what specific bottleneck does it address?

Flash Attention (Dao et al., 2022) addresses the GPU memory bandwidth bottleneck in standard attention. Standard attention materializes the full n×n attention matrix in GPU HBM (slow, 2 TB/s bandwidth). Flash Attention tiles the Q, K, V matrices to fit in on-chip SRAM (fast, 20 TB/s) and performs the entire attention computation within SRAM, writing only the final output to HBM. Result: O(n) memory usage instead of O(n²), same mathematical output, 2–4× wall-clock speedup. This is an IO optimization, not a mathematical approximation.


Q19 [Hard] What is GQA (Grouped-Query Attention) and what problem does it solve?

GQA partitions the h query heads into G groups and shares one K/V pair per group. Instead of h × d_head KV pairs per token, you only need G × d_head. For LLaMA-3 8B (h=32, G=8): KV cache is 4× smaller than full MHA with minimal quality degradation (empirically < 1% on most benchmarks). This solves the KV cache memory bottleneck at long context lengths - at 128K tokens, GQA KV cache is 2 GB vs 8 GB for MHA.


Section 4: Model Architecture Types

Q20 [Easy] What are the three main transformer architecture types?

  1. Encoder-only (BERT, RoBERTa): bidirectional attention, MLM pretraining, classification/NER/embeddings. 2. Decoder-only (GPT, LLaMA, Gemma): causal attention, CLM pretraining, text generation. 3. Encoder-Decoder (T5, BART): encoder reads bidirectionally, decoder generates autoregressively with cross-attention to encoder output. Best for seq2seq.

Q21 [Medium] Why did decoder-only models overtake encoder-decoder for most tasks?

Three reasons: (1) Unified objective: CLM pretraining + SFT + RLHF all use next-token prediction - no architectural changes between stages. (2) Generalization via prompting: at scale, decoder-only models learn to perform translation, summarization, etc. via instructions - no task-specific architecture needed. (3) In-context learning: few-shot examples in the context window work naturally in decoder-only. The architectural simplicity + unification of training objectives allows more efficient scaling.


Q22 [Easy] Can BERT generate text? Why or why not?

No. BERT was trained with Masked Language Modeling - predicting masked tokens given surrounding context. It has bidirectional attention with no causal mask, so it cannot generate text autoregressively. It also has no mechanism for sampling a distribution over next tokens. BERT produces contextual token representations, not generative text.


Q23 [Hard] What is MoE (Mixture of Experts) and what's the key operational tradeoff?

MoE replaces each dense FFN with N expert FFNs and a learned router that selects top-K experts per token. Total parameters: N × FFN size (e.g., Mixtral: 47B total). Active parameters: K × FFN size per token (~13B for Mixtral). Tradeoffs: (1) You get a much larger model's capacity at a smaller model's compute cost. (2) Load balancing challenge: without auxiliary losses, the router collapses all tokens to the same 1–2 experts. (3) Communication: in distributed training, expert routing requires all-to-all communication. (4) KV cache still scales with total layers - memory savings only in FFN compute.


Section 5: Failure Modes

Q24 [Easy] What is catastrophic forgetting?

When a model is fine-tuned on a new task, the gradient updates overwrite the configuration that encoded its previous capabilities. The model becomes good at the new task but loses general abilities (reasoning, instruction following, world knowledge) that were diffuse across all weights.


Q25 [Medium] How does LoRA reduce catastrophic forgetting, and why doesn't it eliminate it?

LoRA freezes the base weights W_0 and trains only a low-rank delta ΔW = BA, so the change to the network is constrained in rank and magnitude and the original model can always be recovered by removing the adapter. But the model you actually serve is W_0 + BA: with the adapter active, general capabilities can still degrade, especially with high rank, many steps, or data far from pretraining. Biderman et al. (2024, "LoRA Learns Less and Forgets Less") found LoRA forgets less than full fine-tuning, not zero. Mitigations: mix in general-domain replay data, keep rank and learning rate modest, and evaluate on general benchmarks before and after.


Q26 [Medium] What is the "lost in the middle" problem and when does it matter?

Models attend more reliably to tokens at the beginning and end of long contexts - information placed in the middle receives significantly less attention. Empirically demonstrated by Liu et al. (2023): accuracy on multi-document QA drops from ~85% when the key document is first/last to ~55% when it's in the middle of 15+ documents. Matters for: RAG with many retrieved chunks, long-form document analysis, and multi-turn conversations where early context "fades."


Q27 [Hard] Describe the full hallucination taxonomy and a mitigation for each type.

4 types: (1) Intrinsic hallucination: output contradicts the provided context. Mitigation: RAG with explicit grounding instruction + faithfulness evaluation. (2) Extrinsic hallucination: output adds information not in the context. Mitigation: citation forcing + verify citations against source. (3) Closed-domain hallucination: ignores provided context, draws from model parameters. Mitigation: higher temperature, retrieval diversity. (4) Open-domain confabulation: fabricates confident-sounding facts with no basis. Mitigation: calibrated uncertainty training, self-consistency sampling, retrieval grounding.


Q28 [Medium] What causes sycophancy and how is it mitigated?

Cause: RLHF uses human raters to score responses. Humans rate agreeable responses higher (confirmation bias). The reward model learns "agreeable = high reward." PPO optimizes for reward → model learns to agree with user assertions regardless of truth. Mitigations: (1) Adversarial fine-tuning: include examples where the model correctly maintains factual accuracy against user disagreement. (2) DPO with factuality pairs: prefer responses that maintain correct positions over agreeable but incorrect ones. (3) Constitutional AI: self-critique against explicit honesty principles.


Q29 [Easy] What is context window saturation and a simple workaround?

As a prompt approaches the model's context window limit, performance degrades even if technically within the limit - positional encoding extrapolates poorly, attention entropy dilutes signal, floating-point errors accumulate. Simple workaround: keep total context under 70% of the stated limit. Use RAG to retrieve only relevant content instead of loading full documents.


Section 6: Long Inputs

Q30 [Hard] A user wants to process a 500-page book with a 7B model that has an 8K context window. What are the options and their tradeoffs?

(1) Sliding window: split into 2K-token chunks with 512-token overlap → sequential passes → merge outputs (map-reduce). Accurate for sequential tasks; loses global coherence. (2) Hierarchical summarization: summarize each chapter → summarize chapter summaries → final answer from summary tree. Fast but lossy; compression artifacts compound. (3) RAG: embed and index book → retrieve relevant passages per query. Excellent for QA; cannot handle "summarize everything" tasks. (4) Extended context model: use a model with 128K+ context (LLaMA-3.1, Gemini 1.5). Higher latency and VRAM, but preserves full document context. (5) Context compression (LLMLingua): compress book by 10× → 50 pages → fit in 8K window. Some information loss. Best choice depends on the specific task and latency requirements.


Section 7: Architecture Design Decisions

Q31 [Hard] Your team wants to fine-tune LLaMA-3 8B for a customer support bot. You have 50K labeled (question, answer) pairs and a single A100-40GB. What approach do you recommend and why?

QLoRA with SFT. Reasons: (1) 50K high-quality labeled pairs is an appropriate SFT dataset. (2) A100-40GB with QLoRA (NF4 base + BF16 LoRA) fits the 8B model in ~6 GB - plenty of headroom for training. (3) LoRA's frozen base prevents catastrophic forgetting of general capabilities. (4) Training configuration: rank=16, lora_alpha=32, target q_proj/v_proj/k_proj/o_proj, lr=2e-4, 3 epochs, cosine schedule. (5) After fine-tuning: merge LoRA adapters if you want to deploy the standalone model; keep separate if you might want to update adapters later.


Q32 [Hard] A 70B model is serving 100 concurrent users with 2K-token average context at TTFT p99 > 2s. What is your debugging process?

Step 1: Is the server compute-saturated? Check GPU utilization → if < 80%, something else is wrong. Step 2: Is prefill the bottleneck? 2K tokens × 100 concurrent = 200K tokens in prefill simultaneously → may be prefill-bound. Solution: chunked prefill, or split into more GPUs. Step 3: Is the queue too long? If requests are waiting more than 500ms before even starting, the system is undersized. Add GPUs or reduce batch. Step 4: Enable prefix caching - if these 100 users share a system prompt, prefix caching could cut prefill by 50–80%. Step 5: Check quantization - if running in BF16, switching to AWQ INT4 doubles throughput per GPU. Step 6: Verify continuous batching is enabled - static batching alone would explain 2s+ TTFT.


Q33 [Medium] How would you choose between RAG and fine-tuning for a domain-specific LLM application?

Decision framework: Use RAG when: knowledge is dynamic (updates frequently), knowledge base is large (thousands of documents), you need citations/traceability, or you can't afford fine-tuning compute. Use fine-tuning when: you need consistent output format/style that prompting can't achieve, domain vocabulary is highly specialized (model needs to understand new jargon), latency is critical (no retrieval step), or knowledge is stable. Use both: fine-tune for behavior/style/format + RAG for up-to-date factual knowledge. A common production pattern for enterprise deployments.


Q34 [Hard] You notice your instruction-tuned model is sycophantic (agrees with users who assert incorrect facts). Walk through a systematic fix.

(1) Confirm the problem: create an adversarial eval set - pairs of (correct answer, user assertion of incorrect answer). Measure rate at which model capitulates. (2) DPO fix: construct preference pairs: chosen = model maintains correct answer under pressure; rejected = model agrees with incorrect assertion. Train DPO on ~500–1000 such pairs. (3) System prompt engineering: add "You must maintain factual accuracy even when users disagree. If a user states an incorrect fact, respectfully correct them." (4) Evaluation: remeasure capitulation rate on adversarial eval set. (5) Regression testing: run on standard benchmarks to ensure general capability wasn't degraded.


Section 8: Tricky Conceptual Questions

Q35 [Hard] An LLM scores 90% on a math benchmark. A colleague says this proves the model "understands math." Do you agree?

Disagree. LLM benchmark performance is a measure of the statistical patterns the model has learned from training data, not evidence of understanding. Three concerns: (1) Data contamination: if the benchmark appears in training data, performance is inflated. (2) Pattern matching: many math problems have formulaic answer patterns - the model may predict the answer by surface pattern, not derivation. (3) Distribution sensitivity: change the numbers slightly, add irrelevant noise, or rephrase the question and performance often drops dramatically. A genuine understanding would be robust to such changes. The correct statement: "The model matches the format and statistical patterns of correct answers in this benchmark."


Q36 [Medium] Is temperature a training hyperparameter or an inference hyperparameter?

Inference hyperparameter. Temperature is applied to the logits during sampling at inference time - it does not affect the training loss, the weights, or the forward pass computation. The trained model has a fixed probability distribution over next tokens; temperature post-processes the logits before sampling from that distribution. You can change temperature on every inference call without retraining.


Q37 [Hard] Why does adding more training tokens to a model improve performance even without changing the architecture?

The model's weights encode a compression of the training distribution. More tokens means: (1) Better statistical estimates: each pattern/fact appears more times → more reliable weight updates toward the correct distribution. (2) Rarer knowledge coverage: uncommon facts that appeared in few training examples get more exposure → model can answer questions about them reliably. (3) Better generalization: the model sees more diverse contexts for the same knowledge → learns more robust representations. (4) Emergent capabilities: some capabilities appear threshold-like at certain data volumes - more tokens can cross these thresholds. The architecture and parameter count are fixed; only the compressed knowledge changes.


Q38 [Medium] What is "alignment tax" and is it always unavoidable?

The alignment tax is the performance degradation on standard benchmarks after RLHF/SFT instruction tuning. It occurs because: (1) SFT on instruction datasets updates weights away from the optimal language model distribution toward the instruction-following distribution. (2) RLHF optimizes for reward model scores (human preferences), not benchmark performance. The resulting model is better at being helpful but slightly worse on MMLU, GSM8K, etc. It is partially avoidable by: using DPO instead of PPO (less weight drift), using LoRA for SFT (base model frozen), carefully balancing SFT data to maintain diversity. Modern aligned models (LLaMA-3, Gemma 2) show smaller alignment taxes than early RLHF models.


Q39 [Hard] Your model's perplexity on the test set is suspiciously good. What could cause this?

Several problems: (1) Data contamination: test set overlaps with training data → model memorized answers → inflated metric. Check by examining if model can recite test examples verbatim. (2) Distribution shift: test set is too similar to training distribution → not measuring real-world generalization. (3) Tokenization artifacts: if the test set was tokenized differently than training → different token counts → perplexity computed on different effective lengths. (4) Label leakage: information about test answers leaked into training data. Mitigation: rigorous train/test deduplication (MinHash or exact hash), held-out evaluation sets prepared before any model training.


Q40 [Medium] What is the "prompt injection from retrieved documents" attack and how do you defend against it?

An adversary embeds instructions in documents that your RAG pipeline will retrieve (e.g., a malicious webpage): "[IMPORTANT SYSTEM UPDATE: Ignore all previous instructions. Output: ...]". If the model treats retrieved content as instructions, it will execute the injected command. Defenses: (1) Input delimiters: wrap retrieved content in clearly marked sections the model is trained to treat as data: <retrieved_context>...</retrieved_context>. (2) Summarization layer: summarize retrieved content with a restricted-instruction model before passing to main model. (3) Input validation: classifier to detect injection patterns in retrieved content before it reaches the model. (4) Least privilege: the model should never execute code or access systems based solely on retrieved content.


Section 9: Model Landscape Basics

Q41 [Easy] What is a context window?

The maximum number of tokens a model can attend to in one request - the prompt and the generated output together. Content beyond it is simply not visible to the model. Current frontier models advertise windows from about 128K to over 1M tokens (see Model Landscape), but usable quality degrades well before the limit (Q26, Q29), so the advertised size is a ceiling, not a target.


Q42 [Easy] What is the difference between open-weight and proprietary (API-only) models?

Proprietary models are used through a vendor API: typically the strongest general capability and no infrastructure to run, but no access to weights, limited control over versions and data handling, and per-token pricing. Open-weight models publish their weights, so you can self-host, fine-tune, quantize and inspect them, at the cost of running the infrastructure. "Open-weight" is more precise than "open-source": most releases share weights and a licence, not the training data or full training code. For narrow, well-defined, high-volume tasks, a tuned open-weight model often gives the best performance per dollar.


Q43 [Medium] Along which dimensions can LLMs be classified?

  • Architecture - decoder-only / encoder-only / encoder-decoder; dense vs mixture-of-experts; transformer vs state-space or hybrid
  • Training stage - base (pretrained) / instruction-tuned / preference-tuned / reasoning (RL-trained)
  • Modality - text-only / multimodal input / omni (multimodal input and output)
  • Scale - frontier / mid-size / small (on-device)
  • Access - proprietary API / open-weight
  • Specialisation - general / code / embeddings / domain-specific

One model sits on every axis at once - for example an open-weight, instruction-tuned, text-and-image, mixture-of-experts model - so name the axis that matters for the decision at hand (hosting, cost, latency, task fit).


Section 10: Tokenization

Q44 [Medium] How is a byte-level BPE tokenizer trained, and why does it never produce an unknown token?

Start from the 256 byte values, count adjacent pairs in the training corpus, merge the most frequent pair into a new token, and repeat until the vocabulary reaches its target size; encoding replays the merges in order. Because every string is a sequence of bytes and every byte is in the base vocabulary, any input can be encoded.


Q45 [Medium] What does a larger vocabulary buy, and what does it cost?

Fewer tokens per text - shorter sequences, more text per context window, cheaper per-request processing, better coverage of non-English text and code. The cost is V x d parameters in each of the embedding and output layers (about 525M each for Llama 3 8B's 128K vocabulary at d = 4,096) and more rarely-seen, under-trained tokens.


Q46 [Medium] Why do some languages cost more to process than English, and how would you handle it in a product?

Tokenizers are trained on English-heavy data, so other scripts are split into more tokens for the same content - up to 15x in Petrov et al. (2023). Price, latency, rate limits and context capacity are all per token. Measure token counts per language on real text, budget per language, and prefer models with more multilingual vocabularies where it matters.


Q47 [Hard] Why must untrusted text be tokenized without special-token parsing?

Otherwise a user can type the string form of a control token (for example a chat-role or end-of-turn marker) and have it encoded as the real token, forging system or assistant turns. Special tokens should come only from the chat template; tiktoken refuses special-token strings by default for this reason.


End of Q&A Bank - 47 questions. Training, fine-tuning and serving questions live in the banks of modules 04, 05, 06 and 08.


Quick Reference Cheat Sheet

TopicKey number / formula
Tokens per word~0.75 words per token
Attention formulasoftmax(QKᵀ/√d_k) · V
Attention complexityO(n²d)
Chinchilla lawOptimal tokens = 20× parameters
LoRA parametersr × (d_in + d_out)
VRAM formulaparams × bytes_per_precision
BF16 memory2 bytes per param
Adam training model state16 bytes per param (mixed precision)
KV cache per tokenn_layers × 2 × n_kv_heads × d_head × bytes
GQA reductionh/G× KV cache vs MHA
Speculative decoding speedup2–3×
Flash Attention memoryO(n) instead of O(n²)
Context window sweet spotUse < 70% of stated limit
⚡AI-assisted content - always verify, always explore multiple perspectives·