02 - PyTorch Fundamentals
Tensors, autograd, data loading, the hand-written training loop, checkpointing, mixed precision and GPU-memory debugging - then the tools used to train LLMs (fused attention, torch.compile, bf16, DDP and FSDP2) and a small modern GPT built from scratch.
Learning objectives 8-10 hours (notes + labs)
By the end of this module you will be able to:- Build and debug a model in raw PyTorch - tensors, autograd, Dataset/DataLoader, the training and validation loop
- Checkpoint and resume training exactly, and train in mixed precision (bf16, or fp16 with GradScaler)
- Diagnose device, shape and out-of-memory errors and budget GPU memory for a training job
- Use the LLM toolkit - SDPA/FlexAttention, torch.compile, torchrun with DDP and FSDP2, safetensors and distributed checkpoints
- Build and train a small modern decoder (RMSNorm, RoPE, GQA, SwiGLU) from scratch
Prerequisites
- Python (classes, list comprehensions), NumPy, and basic calculus
Chapter Map
| # | File | Topic | Difficulty |
|---|---|---|---|
| 1 | Tensors & Autograd | Tensor basics, requires_grad, computation graph, .backward(), gradient accumulation | Beginner |
| 2 | Dataset & DataLoader | Custom Dataset, DataLoader batching/shuffling/num_workers, collation | Beginner |
| 3 | Training Loop From Scratch | Hand-written epoch/batch loop, model.train()/model.eval(), validation loop | Intermediate |
| 4 | Checkpointing & Mixed Precision | state_dict() save/load, resuming training, autocast, GradScaler | Intermediate |
| 5 | Debugging & GPU Memory | .to(device), CUDA OOM errors, shape-mismatch debugging, torch.cuda.memory_summary() | Intermediate |
| 6 | PyTorch for LLMs | SDPA/FlexAttention, torch.compile, bf16, torchrun, DDP vs FSDP2, safetensors and distributed checkpoints, profiling and MFU | Advanced |
| 7 | Q&A Review Bank | 29 Q&A pairs across all topics | All levels |
Code Labs
| Lab | What you build |
|---|---|
| Raw PyTorch Classifier | An MNIST CNN with the hand-written loop, checkpointing and AMP - the warm-up |
| GPT From Scratch | A ~9.5M-parameter modern decoder trained on Tiny Shakespeare - the bridge to Pretraining at Scale |
Recommended Learning Paths
Path A: Beginner - First PyTorch Model
- Tensors & Autograd - understand tensors and how gradients flow
- Dataset & DataLoader - learn to feed data into a model
- Training Loop From Scratch - write your first end-to-end training loop
- Raw PyTorch Classifier - apply everything in a runnable CNN
Path B: Interview Preparation (Accelerated)
- Tensors & Autograd - gradient accumulation and
.backward()questions are very common - Training Loop From Scratch - train/eval mode, why
zero_grad()matters - Checkpointing & Mixed Precision - AMP questions come up in production-focused interviews
- Q&A Review Bank - drill all 29 questions
Path C: Production/Debugging Focus (Advanced)
- Debugging & GPU Memory - CUDA OOM triage and shape-mismatch fixes
- Checkpointing & Mixed Precision - resuming long-running training jobs
- Raw PyTorch Classifier - a complete, checkpointed, AMP-enabled training run
Path D: Toward LLM Training (Advanced)
- PyTorch for LLMs - fused attention, compile, bf16, FSDP2, checkpoints
- GPT From Scratch - build and train a modern decoder end to end
- Continue to Pretraining at Scale
Resources
- Q&A Review Bank - 29 Q&A pairs in this module
- Module quiz - every Check Yourself question in this module
- Raw PyTorch Classifier Code Lab - hands-on CNN trained end-to-end
Key Cross-References
- Pretraining objectives and how LLMs are trained at scale → Training & Pretraining
- LoRA/QLoRA fine-tuning built on the same
autograd+ optimizer mechanics covered here → Fine-Tuning - VRAM estimation and mixed-precision tradeoffs at model scale → GPU & Hardware
- Hands-on fine-tuning that builds directly on these mechanics → Fine-Tuning Lab
Section Appendix
Summary & Key Terms - a quick recap of this section and its essential vocabulary.
Next Topic
Previous: 02 - Python & Systems · Next: 03 - Math for ML