Contents
Map

03 · Math for ML

Appendix - Summary & Key Terms

View as:

Appendix - Math for ML

What We Learned

  • Track matrix shapes to explain attention, parameter counts and compute.
  • Probability, softmax and cross-entropy connect next-token prediction to training loss.
  • The chain rule yields backpropagation; learning rate and optimizer choices affect stability.
  • Use paired comparisons and confidence intervals to distinguish improvements from noise.
  • Account for clustered items, repeated trials, judge agreement and multiple comparisons.

Key Acronyms, Concepts & Jargon

TermShort meaning
Dot product / cosine similarityVector alignment / alignment after removing magnitude differences.
SVDSingular Value Decomposition: factors a matrix into rotations and singular-value scaling.
RankNumber of independent directions represented by a matrix.
FLOPsFloating-Point Operations: a count of numerical work.
SoftmaxConverts logits into probabilities that sum to one.
NLL / cross-entropyNegative Log-Likelihood / prediction loss; equivalent for observed target labels.
KL divergenceKullback-Leibler divergence: an asymmetric measure of distribution mismatch.
PerplexityExponential of average token NLL in nats; depends on the tokenizer.
Gradient / backpropagationLoss derivatives / efficient chain-rule computation through a model.
SGD / AdamWStochastic Gradient Descent / adaptive optimizer with decoupled weight decay.
CI / SEConfidence Interval / Standard Error: uncertainty range / estimated sampling variability.
Bootstrap / paired testResampling to estimate uncertainty / comparing systems on the same items.
pass@k / pass^kAt least one success in k attempts / success on every one of k trials.
Cohen's kappaJudge agreement corrected for agreement expected by chance.

Back to section overview

⚡AI-assisted content - always verify, always explore multiple perspectives·