Contents
Map

04 · Pretraining at Scale

Overview

View as:

04 - Pretraining at Scale

Pretraining turns trillions of tokens of text into a base model by predicting the next token, over and over, on thousands of GPUs. This module covers what goes into that run - the data pipeline, the objective, how compute is split across GPUs, and the memory and hardware budget that decides what is feasible.

Learning objectives 8-10 hours
By the end of this module you will be able to:
  • Walk through a pretraining data pipeline - filtering, deduplication, quality scoring, mixing - and explain why each stage exists
  • Apply scaling laws to size a model and its token budget for a compute budget
  • Estimate training memory per parameter and explain how data, tensor, pipeline and ZeRO/FSDP parallelism divide it
  • Explain the interconnect and precision choices that bound training throughput, and use the roofline model
  • Choose a parallelism layout and precision (BF16/FP8) for a run, and estimate its MFU and cost
  • Diagnose training instability and choose schedules and optimizers (WSD, μP, Muon)
  • Profile a training step and fix GPU bottlenecks - data loading, launch overhead, unfused memory-bound kernels, communication
Prerequisites

Chapter Map

#NoteTopicLevel
1Pretraining OverviewThe big picture of a pretraining run, tokenizer choices for pretraining (data budgets, vocabulary size), training objectives (CLM, MLM, span corruption), and a map of the chapters belowIntermediate
2GPU Memory & HardwareMemory estimation (16 B/param), precision formats, ZeRO/FSDP sharding arithmetic, memory hierarchy and interconnectsIntermediate
3Data Curation & MixturesDedup (MinHash LSH), model-based quality filtering (FineWeb-Edu, DCLM), mixtures, annealing, mid-training, long-context stages, synthetic dataAdvanced
4Scaling Laws & Compute BudgetsC ≈ 6ND, Chinchilla fit, inference-aware and data-constrained scaling, test-time computeAdvanced
5Distributed Training at Scale5D parallelism layout, pipeline bubbles, ring attention, expert parallelism, FP8 training, MFU, fault toleranceAdvanced
6Training Stability & OptimizersLoss spikes, QK-norm, z-loss, WSD schedules, μP, Muon and MuonClipAdvanced
7Accelerators & InterconnectsH100 → B200 / GB200 NVL72, AMD Instinct, TPU Ironwood, roofline model, FP8/MXFP4/NVFP4, scale-up vs scale-outAdvanced
8CUDA Concepts & GPU ProfilingThreads, warps and SMs, memory hierarchy, kernel fusion and Triton, timing async GPU code, torch.profiler / Nsight Systems / Nsight Compute workflowAdvanced
9Q&A Review Bank35 questions across data, scaling, distributed training, hardware, stability and profilingAll levels

Suggested order: 1 → 2 → 4 → 3 → 5 → 6 → 7 → 8, then the bank. Pair chapter 5 with the GPT From Scratch lab to compute MFU on a real run, and chapter 8 with the Profile and Fuse lab.

Code Lab

  • Profile and Fuse - profile a GPT training step with torch.profiler, then fuse a memory-bound SwiGLU + RMSNorm tail with torch.compile and a Triton kernel, with bootstrap confidence intervals on the speedup. The CPU path runs on any laptop; the Triton path needs a CUDA GPU.

Review

Section Appendix

Summary & Key Terms - a quick recap of this section and its essential vocabulary.


Next Topic

Previous: 03 - Math for ML · Next: 05 - Post-Training & Reasoning

⚡AI-assisted content - always verify, always explore multiple perspectives·