Contents
Map

04 · Pretraining at Scale

Appendix - Summary & Key Terms

View as:

Appendix - Pretraining at Scale

What We Learned

  • Pretraining learns broad patterns; curate, deduplicate, decontaminate and mix data before training.
  • Budget model size, tokens and compute together; deployment costs can justify extra training.
  • GPU memory includes weights, gradients, optimizer state and activations.
  • Combine parallelism strategies to fit the model and match communication to hardware links.
  • Stabilize runs with schedules, clipping and checkpoints; profile compute, memory and communication.

Key Acronyms, Concepts & Jargon

TermShort meaning
CLM / MLMCausal / Masked Language Modeling: predict the next token / selected hidden tokens.
Deduplication / decontaminationRemove repeated data / evaluation overlap from training data.
MinHash / LSHSimilarity sketch / Locality-Sensitive Hashing: find likely near-duplicate documents.
Scaling lawEmpirical relationship between loss, model size, data and compute.
DP / TP / PPData / Tensor / Pipeline Parallelism: split batches / layer computations / model stages.
CP / EPContext / Expert Parallelism: split sequences / distribute MoE experts.
ZeRO / FSDPZero Redundancy Optimizer / Fully Sharded Data Parallel: shard training state.
MFUModel FLOPs Utilization: useful model compute relative to accelerator peak compute.
HBM / SRAMHigh Bandwidth Memory / Static RAM: accelerator main memory / fast on-chip memory.
PCIe / NVLinkPeripheral Component Interconnect Express / NVIDIA GPU interconnect: device communication links.
BF16 / FP8Brain Floating Point 16 / 8-bit Floating Point: lower-precision numerical formats.
CUDA / kernel fusionNVIDIA's GPU computing platform / combining operations to cut launches and memory traffic.
Roofline modelEstimates performance limits from compute capacity and memory bandwidth.
Gradient clipping / warmupCaps gradient size / gradually raises the initial learning rate.

Back to section overview

⚡AI-assisted content - always verify, always explore multiple perspectives·