| Tensor / autograd | Multidimensional array / automatic differentiation through recorded operations. |
| Batch / epoch | Examples in one step / one pass through the training data. |
| Dataset / DataLoader | Sample-access interface / batching, shuffling and loading machinery. |
| Gradient accumulation | Adds gradients across micro-batches before one optimizer update. |
| train() / eval() | Selects training / evaluation layer behavior; eval does not disable gradients. |
| Checkpoint | Saved weights and run state, including optimizer, scheduler and randomness for resumption. |
| AMP / BF16 | Automatic Mixed Precision / Brain Floating Point 16: mixed dtypes / a wide-range 16-bit format. |
| OOM / VRAM | Out Of Memory / Video Random Access Memory: allocation failure / GPU memory. |
| Activation checkpointing | Recomputes intermediate activations during backward to save memory. |
| SDPA | Scaled Dot-Product Attention: the attention operation, with optimized PyTorch backends. |
| DDP / FSDP | Distributed Data Parallel / Fully Sharded Data Parallel: replicate / shard model state across workers. |
| torch.compile | Compiles model execution to reduce overhead and optimize operations. |
| GPT / CNN | Generative Pre-trained Transformer / Convolutional Neural Network: text decoder / convolution-based model. |