Contents
Map

02 · Prog Langs

Overview

View as:

02 - PyTorch Fundamentals

Tensors, autograd, data loading, the hand-written training loop, checkpointing, mixed precision and GPU-memory debugging - then the tools used to train LLMs (fused attention, torch.compile, bf16, DDP and FSDP2) and a small modern GPT built from scratch.

Learning objectives 8-10 hours (notes + labs)
By the end of this module you will be able to:
  • Build and debug a model in raw PyTorch - tensors, autograd, Dataset/DataLoader, the training and validation loop
  • Checkpoint and resume training exactly, and train in mixed precision (bf16, or fp16 with GradScaler)
  • Diagnose device, shape and out-of-memory errors and budget GPU memory for a training job
  • Use the LLM toolkit - SDPA/FlexAttention, torch.compile, torchrun with DDP and FSDP2, safetensors and distributed checkpoints
  • Build and train a small modern decoder (RMSNorm, RoPE, GQA, SwiGLU) from scratch
Prerequisites
  • Python (classes, list comprehensions), NumPy, and basic calculus

Chapter Map

#FileTopicDifficulty
1Tensors & AutogradTensor basics, requires_grad, computation graph, .backward(), gradient accumulationBeginner
2Dataset & DataLoaderCustom Dataset, DataLoader batching/shuffling/num_workers, collationBeginner
3Training Loop From ScratchHand-written epoch/batch loop, model.train()/model.eval(), validation loopIntermediate
4Checkpointing & Mixed Precisionstate_dict() save/load, resuming training, autocast, GradScalerIntermediate
5Debugging & GPU Memory.to(device), CUDA OOM errors, shape-mismatch debugging, torch.cuda.memory_summary()Intermediate
6PyTorch for LLMsSDPA/FlexAttention, torch.compile, bf16, torchrun, DDP vs FSDP2, safetensors and distributed checkpoints, profiling and MFUAdvanced
7Q&A Review Bank29 Q&A pairs across all topicsAll levels

Code Labs

LabWhat you build
Raw PyTorch ClassifierAn MNIST CNN with the hand-written loop, checkpointing and AMP - the warm-up
GPT From ScratchA ~9.5M-parameter modern decoder trained on Tiny Shakespeare - the bridge to Pretraining at Scale

Path A: Beginner - First PyTorch Model

  1. Tensors & Autograd - understand tensors and how gradients flow
  2. Dataset & DataLoader - learn to feed data into a model
  3. Training Loop From Scratch - write your first end-to-end training loop
  4. Raw PyTorch Classifier - apply everything in a runnable CNN

Path B: Interview Preparation (Accelerated)

  1. Tensors & Autograd - gradient accumulation and .backward() questions are very common
  2. Training Loop From Scratch - train/eval mode, why zero_grad() matters
  3. Checkpointing & Mixed Precision - AMP questions come up in production-focused interviews
  4. Q&A Review Bank - drill all 29 questions

Path C: Production/Debugging Focus (Advanced)

  1. Debugging & GPU Memory - CUDA OOM triage and shape-mismatch fixes
  2. Checkpointing & Mixed Precision - resuming long-running training jobs
  3. Raw PyTorch Classifier - a complete, checkpointed, AMP-enabled training run

Path D: Toward LLM Training (Advanced)

  1. PyTorch for LLMs - fused attention, compile, bf16, FSDP2, checkpoints
  2. GPT From Scratch - build and train a modern decoder end to end
  3. Continue to Pretraining at Scale

Resources

Key Cross-References

  • Pretraining objectives and how LLMs are trained at scale → Training & Pretraining
  • LoRA/QLoRA fine-tuning built on the same autograd + optimizer mechanics covered here → Fine-Tuning
  • VRAM estimation and mixed-precision tradeoffs at model scale → GPU & Hardware
  • Hands-on fine-tuning that builds directly on these mechanics → Fine-Tuning Lab

Section Appendix

Summary & Key Terms - a quick recap of this section and its essential vocabulary.


Next Topic

Previous: 02 - Python & Systems · Next: 03 - Math for ML

⚡AI-assisted content - always verify, always explore multiple perspectives·