Contents
Map

03 ยท Math for ML

Overview

View as:

03 - Math for ML

The math that LLM engineering actually uses, taught through where it shows up: linear algebra for tensor shapes, attention, parameter counts and LoRA; probability and information theory for the next-token distribution, sampling, the training loss, perplexity and KL penalties; calculus and optimization for backpropagation, Adam and learning-rate schedules; and statistics for comparing models without fooling yourself. No proofs for their own sake - every idea is tied to a model, a training run or an eval.

Learning objectives 5-6 hours (notes + lab)
By the end of this module you will be able to:
  • Track tensor shapes through a transformer layer and count its parameters and FLOPs
  • Explain the training loss as maximum likelihood, and convert between loss, perplexity and bits
  • Derive gradients with the chain rule, and explain how Adam, weight decay and learning-rate schedules shape training
  • Report an eval score with a confidence interval and compare two models with a paired test
  • Implement attention and its backward pass by hand and verify the gradients
Prerequisites
  • High-school algebra and some Python - the notes introduce everything else
  • PyTorch Fundamentals is helpful for the lab

Where This Module Fits

flowchart LR
    LA["๐Ÿ“ 01 Linear algebra<br/>shapes, rank, SVD"] --> PR["๐ŸŽฒ 02 Probability &<br/>information theory"]
    PR --> CA["๐Ÿ“‰ 03 Calculus &<br/>optimization"]
    PR --> ST["๐Ÿ“Š 04 Statistics<br/>for evaluation"]
    CA --> LAB["๐Ÿงช Lab: attention &<br/>backprop by hand"]
    LA --> LAB
    LAB --> NEXT["๐Ÿ‹๏ธ Pretraining ยท Post-training ยท Evaluation"]
    ST --> NEXT

    style LA fill:#d8dfe8,stroke:#b0bac8
    style PR fill:#e8e0d4,stroke:#c8b89a
    style CA fill:#dde4dc,stroke:#b0c4b0
    style ST fill:#ddd8e4,stroke:#b8b0c8

Chapter Map

#ChapterYou will learnTime
1Linear Algebra for TransformersDot products and cosine similarity, shapes through a layer, parameter and FLOP counts, rank and LoRA, SVD, norms50 min
2Probability & Information TheoryThe LM as a product of conditionals, softmax and temperature, maximum likelihood, perplexity, entropy, KL, Bayes and base rates50 min
3Calculus & OptimizationChain rule and backprop, key gradients, learning-rate stability, momentum, Adam/AdamW, schedules, exploding and vanishing gradients55 min
4Statistics for EvaluationStandard errors and Wilson intervals, bootstrap, paired tests and McNemar, clustering, multiple comparisons, sample size, pass@k, kappa50 min
5Q&A Review Bank16 questions across the module30 min

Code Lab

  • Attention & Backprop by Hand - one attention head in NumPy, gradients derived by hand and verified against finite differences and PyTorch autograd, plus experiments on โˆšd scaling and temperature. Runs on any laptop in seconds.

How Much Math Do You Need?

For building applications - prompting, RAG, agents - notes 2 and 4 are the most useful: sampling, loss and perplexity, and reading eval results honestly. For training, fine-tuning and serving work, all four notes are the working vocabulary of the papers and code in modules 04-08. Nothing here requires calculus beyond the chain rule or linear algebra beyond matrix multiplication and SVD.

Resources

Key Cross-References

Section Appendix

Summary & Key Terms - a quick recap of this section and its essential vocabulary.


Next Topic

Previous: 02 - PyTorch Fundamentals ยท Next: 04 - Pretraining at Scale

โšกAI-assisted content - always verify, always explore multiple perspectivesยท