Appendix - Math for ML
What We Learned
- Track matrix shapes to explain attention, parameter counts and compute.
- Probability, softmax and cross-entropy connect next-token prediction to training loss.
- The chain rule yields backpropagation; learning rate and optimizer choices affect stability.
- Use paired comparisons and confidence intervals to distinguish improvements from noise.
- Account for clustered items, repeated trials, judge agreement and multiple comparisons.
Key Acronyms, Concepts & Jargon
| Term | Short meaning |
|---|---|
| Dot product / cosine similarity | Vector alignment / alignment after removing magnitude differences. |
| SVD | Singular Value Decomposition: factors a matrix into rotations and singular-value scaling. |
| Rank | Number of independent directions represented by a matrix. |
| FLOPs | Floating-Point Operations: a count of numerical work. |
| Softmax | Converts logits into probabilities that sum to one. |
| NLL / cross-entropy | Negative Log-Likelihood / prediction loss; equivalent for observed target labels. |
| KL divergence | Kullback-Leibler divergence: an asymmetric measure of distribution mismatch. |
| Perplexity | Exponential of average token NLL in nats; depends on the tokenizer. |
| Gradient / backpropagation | Loss derivatives / efficient chain-rule computation through a model. |
| SGD / AdamW | Stochastic Gradient Descent / adaptive optimizer with decoupled weight decay. |
| CI / SE | Confidence Interval / Standard Error: uncertainty range / estimated sampling variability. |
| Bootstrap / paired test | Resampling to estimate uncertainty / comparing systems on the same items. |
| pass@k / pass^k | At least one success in k attempts / success on every one of k trials. |
| Cohen's kappa | Judge agreement corrected for agreement expected by chance. |