03 - Math for ML
The math that LLM engineering actually uses, taught through where it shows up: linear algebra for tensor shapes, attention, parameter counts and LoRA; probability and information theory for the next-token distribution, sampling, the training loss, perplexity and KL penalties; calculus and optimization for backpropagation, Adam and learning-rate schedules; and statistics for comparing models without fooling yourself. No proofs for their own sake - every idea is tied to a model, a training run or an eval.
- Track tensor shapes through a transformer layer and count its parameters and FLOPs
- Explain the training loss as maximum likelihood, and convert between loss, perplexity and bits
- Derive gradients with the chain rule, and explain how Adam, weight decay and learning-rate schedules shape training
- Report an eval score with a confidence interval and compare two models with a paired test
- Implement attention and its backward pass by hand and verify the gradients
- High-school algebra and some Python - the notes introduce everything else
- PyTorch Fundamentals is helpful for the lab
Where This Module Fits
flowchart LR
LA["๐ 01 Linear algebra<br/>shapes, rank, SVD"] --> PR["๐ฒ 02 Probability &<br/>information theory"]
PR --> CA["๐ 03 Calculus &<br/>optimization"]
PR --> ST["๐ 04 Statistics<br/>for evaluation"]
CA --> LAB["๐งช Lab: attention &<br/>backprop by hand"]
LA --> LAB
LAB --> NEXT["๐๏ธ Pretraining ยท Post-training ยท Evaluation"]
ST --> NEXT
style LA fill:#d8dfe8,stroke:#b0bac8
style PR fill:#e8e0d4,stroke:#c8b89a
style CA fill:#dde4dc,stroke:#b0c4b0
style ST fill:#ddd8e4,stroke:#b8b0c8
Chapter Map
| # | Chapter | You will learn | Time |
|---|---|---|---|
| 1 | Linear Algebra for Transformers | Dot products and cosine similarity, shapes through a layer, parameter and FLOP counts, rank and LoRA, SVD, norms | 50 min |
| 2 | Probability & Information Theory | The LM as a product of conditionals, softmax and temperature, maximum likelihood, perplexity, entropy, KL, Bayes and base rates | 50 min |
| 3 | Calculus & Optimization | Chain rule and backprop, key gradients, learning-rate stability, momentum, Adam/AdamW, schedules, exploding and vanishing gradients | 55 min |
| 4 | Statistics for Evaluation | Standard errors and Wilson intervals, bootstrap, paired tests and McNemar, clustering, multiple comparisons, sample size, pass@k, kappa | 50 min |
| 5 | Q&A Review Bank | 16 questions across the module | 30 min |
Code Lab
- Attention & Backprop by Hand - one attention head in NumPy, gradients derived by hand and verified against finite differences and PyTorch autograd, plus experiments on โd scaling and temperature. Runs on any laptop in seconds.
How Much Math Do You Need?
For building applications - prompting, RAG, agents - notes 2 and 4 are the most useful: sampling, loss and perplexity, and reading eval results honestly. For training, fine-tuning and serving work, all four notes are the working vocabulary of the papers and code in modules 04-08. Nothing here requires calculus beyond the chain rule or linear algebra beyond matrix multiplication and SVD.
Resources
- Q&A Review Bank and module quiz
- Deisenroth, Faisal and Ong, Mathematics for Machine Learning - free book covering all four notes in more depth
- Readiness Self-Assessment - this module covers the "linear algebra, probability, optimization" domain
Key Cross-References
- Attention, shapes and FLOPs in context โ Attention Mechanisms, GPT From Scratch
- Optimizers and schedules at scale โ Training Stability & Optimizers
- KL penalties and preference losses โ Preference Optimization
- Statistics applied to real evals โ Building Your Own Evals
Section Appendix
Summary & Key Terms - a quick recap of this section and its essential vocabulary.
Next Topic
Previous: 02 - PyTorch Fundamentals ยท Next: 04 - Pretraining at Scale