04 - Pretraining at Scale
Pretraining turns trillions of tokens of text into a base model by predicting the next token, over and over, on thousands of GPUs. This module covers what goes into that run - the data pipeline, the objective, how compute is split across GPUs, and the memory and hardware budget that decides what is feasible.
Learning objectives 8-10 hours
By the end of this module you will be able to:- Walk through a pretraining data pipeline - filtering, deduplication, quality scoring, mixing - and explain why each stage exists
- Apply scaling laws to size a model and its token budget for a compute budget
- Estimate training memory per parameter and explain how data, tensor, pipeline and ZeRO/FSDP parallelism divide it
- Explain the interconnect and precision choices that bound training throughput, and use the roofline model
- Choose a parallelism layout and precision (BF16/FP8) for a run, and estimate its MFU and cost
- Diagnose training instability and choose schedules and optimizers (WSD, μP, Muon)
- Profile a training step and fix GPU bottlenecks - data loading, launch overhead, unfused memory-bound kernels, communication
Prerequisites
- LLM Foundations - the transformer and attention
- PyTorch Fundamentals - the training loop, mixed precision
- Math for ML - gradients, optimizers and FLOP counting
Chapter Map
| # | Note | Topic | Level |
|---|---|---|---|
| 1 | Pretraining Overview | The big picture of a pretraining run, tokenizer choices for pretraining (data budgets, vocabulary size), training objectives (CLM, MLM, span corruption), and a map of the chapters below | Intermediate |
| 2 | GPU Memory & Hardware | Memory estimation (16 B/param), precision formats, ZeRO/FSDP sharding arithmetic, memory hierarchy and interconnects | Intermediate |
| 3 | Data Curation & Mixtures | Dedup (MinHash LSH), model-based quality filtering (FineWeb-Edu, DCLM), mixtures, annealing, mid-training, long-context stages, synthetic data | Advanced |
| 4 | Scaling Laws & Compute Budgets | C ≈ 6ND, Chinchilla fit, inference-aware and data-constrained scaling, test-time compute | Advanced |
| 5 | Distributed Training at Scale | 5D parallelism layout, pipeline bubbles, ring attention, expert parallelism, FP8 training, MFU, fault tolerance | Advanced |
| 6 | Training Stability & Optimizers | Loss spikes, QK-norm, z-loss, WSD schedules, μP, Muon and MuonClip | Advanced |
| 7 | Accelerators & Interconnects | H100 → B200 / GB200 NVL72, AMD Instinct, TPU Ironwood, roofline model, FP8/MXFP4/NVFP4, scale-up vs scale-out | Advanced |
| 8 | CUDA Concepts & GPU Profiling | Threads, warps and SMs, memory hierarchy, kernel fusion and Triton, timing async GPU code, torch.profiler / Nsight Systems / Nsight Compute workflow | Advanced |
| 9 | Q&A Review Bank | 35 questions across data, scaling, distributed training, hardware, stability and profiling | All levels |
Suggested order: 1 → 2 → 4 → 3 → 5 → 6 → 7 → 8, then the bank. Pair chapter 5 with the GPT From Scratch lab to compute MFU on a real run, and chapter 8 with the Profile and Fuse lab.
Code Lab
- Profile and Fuse - profile a GPT training step with torch.profiler, then fuse a memory-bound SwiGLU + RMSNorm tail with torch.compile and a Triton kernel, with bootstrap confidence intervals on the speedup. The CPU path runs on any laptop; the Triton path needs a CUDA GPU.
Review
- Q&A Review Bank - 35 questions for this module
- Related questions also appear in the LLM Foundations Q&A Review Bank
Section Appendix
Summary & Key Terms - a quick recap of this section and its essential vocabulary.
Next Topic
Previous: 03 - Math for ML · Next: 05 - Post-Training & Reasoning