Contents
Map

Quiz · 03 · Math for ML

24 questions from 5 pages

These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →

Linear Algebra for Transformers

Check yourself
0 / 5 answered
  1. Q has shape (B, h, T, d_h) and K has the same shape. What is the shape of Q @ K.transpose(-2, -1)?
  2. A model has d = 4096 and 32 layers with a classic 4x FFN. Roughly how many parameters are in its transformer layers?
  3. LoRA with rank 8 on a 2048 x 8192 matrix trains how many parameters?
  4. Why does the truncated SVD matter for model compression?
  5. Your embeddings are normalized to unit length. Is ranking by dot product different from ranking by cosine similarity?

Probability & Information Theory

Check yourself
0 / 5 answered
  1. A model's validation loss is 1.5 nats per token. What is its perplexity?
  2. Lowering the sampling temperature from 1.0 to 0.3 does what to the next-token distribution?
  3. Why is minimizing cross-entropy on training data the same as minimizing KL(data ‖ model)?
  4. A judge model agrees with humans 90% of the time on 'bad answer' labels, and flags 10% of good answers as bad. If 5% of production answers are bad, what fraction of flagged answers are actually bad?
  5. What does the KL penalty in RLHF do, and what happens if β is set too low?

Calculus & Optimization

Check yourself
0 / 5 answered
  1. For softmax followed by cross-entropy with a one-hot target y, what is the gradient with respect to the logits?
  2. On the loss L(θ) = 2θ², at what learning rate does gradient descent stop converging?
  3. Roughly how much memory do weights, gradients and AdamW state take for a 13B-parameter model trained in mixed precision?
  4. Why does backpropagation cost only about twice the forward pass, no matter how many parameters there are?
  5. What is the difference between Adam with L2 regularization and AdamW?

Statistics for Evaluation

Check yourself
0 / 5 answered
  1. A model scores 70% on 100 items. Approximately what is the 95% confidence interval using the normal approximation?
  2. Two models are scored on the same 1,000 questions. Which comparison is most appropriate?
  3. Your eval has 50 documents with 10 questions each. Why is the naive standard error over 500 items too small?
  4. With n = 20 samples per problem and c = 2 correct, what is the unbiased pass@1?
  5. You try 30 system-prompt variants and report the best one as a 2-point improvement with p = 0.04. What's wrong?

Attention & Backprop by Hand

Check yourself
0 / 4 answered
  1. Why must the causal mask be applied before the softmax rather than after?
  2. In the lab, X feeds Wq, Wk and Wv. What is dX?
  3. Without the 1/√d_head scaling, what happens to attention as d_head grows, and why does it hurt training?
  4. Why does the finite-difference check use float64?