Contents
Map

02 · Prog Langs

Appendix - Summary & Key Terms

View as:

Appendix - PyTorch Fundamentals

What We Learned

  • Tensors carry shape, dtype and device; autograd accumulates gradients through a computation graph.
  • Dataset fetches samples; DataLoader batches them, with custom collation for variable lengths.
  • Train with forward, loss, backward and parameter updates; validate separately without gradients.
  • Resume with complete checkpoints; mixed precision and memory techniques reduce training cost.
  • Build a classifier and small GPT; profile before applying compilation or distributed training.

Key Acronyms, Concepts & Jargon

TermShort meaning
Tensor / autogradMultidimensional array / automatic differentiation through recorded operations.
Batch / epochExamples in one step / one pass through the training data.
Dataset / DataLoaderSample-access interface / batching, shuffling and loading machinery.
Gradient accumulationAdds gradients across micro-batches before one optimizer update.
train() / eval()Selects training / evaluation layer behavior; eval does not disable gradients.
CheckpointSaved weights and run state, including optimizer, scheduler and randomness for resumption.
AMP / BF16Automatic Mixed Precision / Brain Floating Point 16: mixed dtypes / a wide-range 16-bit format.
OOM / VRAMOut Of Memory / Video Random Access Memory: allocation failure / GPU memory.
Activation checkpointingRecomputes intermediate activations during backward to save memory.
SDPAScaled Dot-Product Attention: the attention operation, with optimized PyTorch backends.
DDP / FSDPDistributed Data Parallel / Fully Sharded Data Parallel: replicate / shard model state across workers.
torch.compileCompiles model execution to reduce overhead and optimize operations.
GPT / CNNGenerative Pre-trained Transformer / Convolutional Neural Network: text decoder / convolution-based model.

Back to section overview

⚡AI-assisted content - always verify, always explore multiple perspectives·