Code Lab 02 - GPT From Scratch
Build and train a small decoder-only transformer in about 250 lines of plain PyTorch, using the same building blocks as current open models - RMSNorm, rotary position embeddings, grouped-query attention through the fused attention kernel, SwiGLU, and weight tying. After this lab, a Llama- or Qwen-style config.json should read like a list of choices you have already made yourself.
← Back to Overview: PyTorch Fundamentals · Concepts: PyTorch for LLMs · Transformer Architecture
- Implement RMSNorm, RoPE, grouped-query attention and a SwiGLU block, and explain each design choice
- Write a language-model training loop with warmup + cosine decay, AdamW with selective weight decay, gradient clipping and bf16 autocast
- Diagnose training from the loss curve, starting from the ln(vocab) sanity check
- Save weights safely with safetensors and sample text with temperature and top-k
What's In This Lab
| Property | Detail |
|---|---|
| Task | Character (byte) level language modeling on Tiny Shakespeare (~1 MB, downloaded on first run) |
| Model | GPT in model.py: 6 layers, 6 query heads, 2 KV heads, d_model 384 - about 9.5M parameters (a tiny 0.4M preset exists for CPU) |
| Training | train.py: random-window batches, AdamW (β₂ = 0.95), warmup + cosine LR, grad-norm clipping, bf16 autocast on GPU, optional torch.compile |
| Output | Train/val loss every 250 steps, checkpoints/model.safetensors, and a 300-byte sample |
| Verified | --preset tiny --steps 400 runs on a laptop CPU in ~10 s: loss falls from 5.57 to ~1.85 and samples look like Shakespeare-shaped text. The default preset needs a GPU for a reasonable run time. |
| Files | 02-GPT-From-Scratch/{model.py, train.py, requirements.txt} |
Architecture
flowchart TD
T["📜 bytes → token ids<br/>(vocab 256)"] --> E["🔤 Embedding"]
E --> N1
subgraph B1["🧱 Decoder block (× n_layer)"]
direction TB
N1["⚖️ RMSNorm"] --> A["🔍 GQA attention<br/>RoPE on q, k · SDPA is_causal"]
A --> R1["➕ residual"]
R1 --> N2["⚖️ RMSNorm"] --> F["🧩 SwiGLU FFN"] --> R2["➕ residual"]
end
R2 --> NF["⚖️ Final RMSNorm"] --> H["🎯 LM head<br/>(tied to embedding)"]
H --> L["📉 Cross-entropy vs next byte"]
style A fill:#d8dfe8,stroke:#b0bac8
style F fill:#dde4dc,stroke:#b0c4b0
style N1 fill:#e8e2d9,stroke:#ccc4b8
style N2 fill:#e8e2d9,stroke:#ccc4b8
style L fill:#e8e0d4,stroke:#c8b89a
Run It
cd 02-Prog-Langs/PyTorch/CodeLabs/02-GPT-From-Scratch
pip install -r requirements.txt
python train.py --preset tiny --steps 400 # CPU smoke test, ~10 seconds
python train.py # default model; use a GPU
python train.py --compile # add torch.compile (first steps are slow while it compiles)
Walkthrough - What to Look At
- The first loss value. With 256 possible bytes and random weights, the model guesses uniformly, so loss starts at ln(256) ≈ 5.55. If your first loss is far from that, initialization or the loss wiring is wrong.
Attention.forward- queries have 6 heads, keys/values only 2.scaled_dot_product_attention(..., enable_gqa=True)shares each KV head across 3 query heads without copying tensors, andis_causal=Trueapplies the mask inside the fused kernel.apply_rope- positions are injected by rotating query/key channel pairs, not by adding a position embedding; nothing positional is learned.build_optimizer- weight decay applies only to matrices; RMSNorm gains are not decayed. β₂ = 0.95 (instead of Adam's default 0.999) is the standard LLM choice because it adapts faster to gradient-scale changes and reduces loss spikes.lr_at- linear warmup then cosine decay. Try the WSD alternative in the exercises.- Checkpoint -
safetensorsstores raw tensors only. A pickled.ptfile can execute arbitrary code when loaded; this is why model hubs moved to safetensors.
Check Yourself
- Your first logged loss is 12.3 with a 256-token vocabulary. What is the most likely problem?
- Why does the model use 2 KV heads for 6 query heads?
- Why is the output projection tied to the input embedding?
Exercises
generate() re-runs the whole context for every new token. Add a KV cache so each step only processes the newest token, and measure the speed-up for 500 generated tokens. Keep RoPE correct: the new token's position is the current cache length.
Hint
Return (k, v) from each Attention call and concatenate along the sequence dimension on the next step.
Hint
With a cache, the single new query attends to all cached keys, so is_causal must be False for those decode steps.
Replace cosine decay with WSD: warm up, hold the peak learning rate for ~80% of training, then decay linearly to min_lr over the final ~20%. Compare the final validation loss with cosine at the same step budget, and explain why WSD makes it cheap to "continue training" from a checkpoint.
Double d_model and n_layer. Predict the parameter count before running (count embedding, attention and SwiGLU matrices), then check with num_params(). Does validation loss improve on Tiny Shakespeare, or does the larger model overfit a 1 MB corpus?
References
- Karpathy, nanoGPT and build-nanogpt - the style this lab follows
- Su et al., RoFormer: Rotary Position Embedding (2021)
- Zhang & Sennrich, Root Mean Square Layer Normalization (2019)
- Shazeer, GLU Variants Improve Transformer (2020)
- Ainslie et al., GQA (2023)
- PyTorch docs:
scaled_dot_product_attention
Last reviewed: 2026-09