01 - LLM Foundations
Every later module - training, serving, prompting, retrieval, agents - rests on one mechanism: a transformer that reads a sequence of tokens and predicts the next one. This module builds that mechanism piece by piece and ends with the ways it predictably fails.
Learning objectives 8-10 hours
By the end of this module you will be able to:- Explain tokenization, next-token prediction and sampling (temperature, top-p) and how they shape model output
- Explain how BPE tokenizers are built and diagnose tokenization-driven failures and costs (numbers, spelling, non-English text)
- Draw a decoder-only transformer block - embeddings, positional encoding (RoPE), attention, normalization, feed-forward, residuals - and say what each part does
- Compute attention by hand and explain multi-head, grouped-query and FlashAttention
- Compare encoder-only, decoder-only, encoder-decoder and mixture-of-experts models and choose one for a task
- Read a current model card - MLA, fine-grained MoE, local-global attention, hybrid layers - and place a model on the landscape
- Recognize the classic failure modes - hallucination, lost in the middle, sycophancy, repetition - and their mitigations
Prerequisites
- Python and basic linear algebra (matrix multiplication, dot products, softmax) - see Math for ML for a refresher
Chapter Map
| # | Note | Topic | Level |
|---|---|---|---|
| 1 | LLM Fundamentals | Tokens, sampling parameters, context windows, model types | Beginner |
| 2 | Tokenization | BPE, WordPiece and Unigram, byte-level tokenizers, vocabulary size, the token tax on non-English text, numbers and glitch tokens, special tokens | Intermediate |
| 3 | Transformer Architecture | Embeddings, positional encoding, Pre-LN and RMSNorm, residuals, SwiGLU | Intermediate |
| 4 | Attention Mechanisms | Q/K/V math, multi-head, causal masking, FlashAttention, GQA | Intermediate |
| 5 | Model Architecture Types | Encoder, decoder, encoder-decoder, mixture of experts, model families | Intermediate |
| 6 | Modern Architectures | MLA, fine-grained MoE, local-global attention, attention sinks, SSM hybrids, native multimodality | Advanced |
| 7 | Failure Modes & Tricky Issues | Forgetting, lost in the middle, hallucination, sycophancy, repetition | Advanced |
| 8 | Model Landscape | How to read and choose among current models; dated snapshot of the major families | Beginner |
| 9 | Q&A Review Bank | 47 questions tagged Easy / Medium / Hard | All levels |
Where the Rest of the Old "LLM Models" Module Went
This module used to cover everything about LLMs. Its later chapters now live where they are taught in depth:
| Topic | Now in |
|---|---|
| Training & pretraining, GPU & hardware | Pretraining at Scale |
| Fine-tuning, RLHF, DPO, GRPO, LoRA | Post-Training & Reasoning |
| KV cache, inference optimization, production deployment | Inference & Serving |
| Prompting mechanics (chat templates, constrained decoding) | Prompt & Context Engineering |
Review
- Module quiz - every Check Yourself question in this module, in course order
Section Appendix
Summary & Key Terms - a quick recap of this section and its essential vocabulary.
Next Topic
Next: 02 - Python & Systems