Quiz · 03 · Math for ML
24 questions from 5 pages
These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
Linear Algebra for Transformers
Check yourself
0 / 5 answered
- Q has shape (B, h, T, d_h) and K has the same shape. What is the shape of Q @ K.transpose(-2, -1)?
- A model has d = 4096 and 32 layers with a classic 4x FFN. Roughly how many parameters are in its transformer layers?
- LoRA with rank 8 on a 2048 x 8192 matrix trains how many parameters?
- Why does the truncated SVD matter for model compression?
- Your embeddings are normalized to unit length. Is ranking by dot product different from ranking by cosine similarity?
Probability & Information Theory
Check yourself
0 / 5 answered
- A model's validation loss is 1.5 nats per token. What is its perplexity?
- Lowering the sampling temperature from 1.0 to 0.3 does what to the next-token distribution?
- Why is minimizing cross-entropy on training data the same as minimizing KL(data ‖ model)?
- A judge model agrees with humans 90% of the time on 'bad answer' labels, and flags 10% of good answers as bad. If 5% of production answers are bad, what fraction of flagged answers are actually bad?
- What does the KL penalty in RLHF do, and what happens if β is set too low?
Calculus & Optimization
Check yourself
0 / 5 answered
- For softmax followed by cross-entropy with a one-hot target y, what is the gradient with respect to the logits?
- On the loss L(θ) = 2θ², at what learning rate does gradient descent stop converging?
- Roughly how much memory do weights, gradients and AdamW state take for a 13B-parameter model trained in mixed precision?
- Why does backpropagation cost only about twice the forward pass, no matter how many parameters there are?
- What is the difference between Adam with L2 regularization and AdamW?
Statistics for Evaluation
Check yourself
0 / 5 answered
- A model scores 70% on 100 items. Approximately what is the 95% confidence interval using the normal approximation?
- Two models are scored on the same 1,000 questions. Which comparison is most appropriate?
- Your eval has 50 documents with 10 questions each. Why is the naive standard error over 500 items too small?
- With n = 20 samples per problem and c = 2 correct, what is the unbiased pass@1?
- You try 30 system-prompt variants and report the best one as a 2-point improvement with p = 0.04. What's wrong?
Attention & Backprop by Hand
Check yourself
0 / 4 answered
- Why must the causal mask be applied before the softmax rather than after?
- In the lab, X feeds Wq, Wk and Wv. What is dX?
- Without the 1/√d_head scaling, what happens to attention as d_head grows, and why does it hurt training?
- Why does the finite-difference check use float64?