Contents
Map

01 · LLM Foundations

Overview

View as:

01 - LLM Foundations

Every later module - training, serving, prompting, retrieval, agents - rests on one mechanism: a transformer that reads a sequence of tokens and predicts the next one. This module builds that mechanism piece by piece and ends with the ways it predictably fails.

Learning objectives 8-10 hours
By the end of this module you will be able to:
  • Explain tokenization, next-token prediction and sampling (temperature, top-p) and how they shape model output
  • Explain how BPE tokenizers are built and diagnose tokenization-driven failures and costs (numbers, spelling, non-English text)
  • Draw a decoder-only transformer block - embeddings, positional encoding (RoPE), attention, normalization, feed-forward, residuals - and say what each part does
  • Compute attention by hand and explain multi-head, grouped-query and FlashAttention
  • Compare encoder-only, decoder-only, encoder-decoder and mixture-of-experts models and choose one for a task
  • Read a current model card - MLA, fine-grained MoE, local-global attention, hybrid layers - and place a model on the landscape
  • Recognize the classic failure modes - hallucination, lost in the middle, sycophancy, repetition - and their mitigations
Prerequisites
  • Python and basic linear algebra (matrix multiplication, dot products, softmax) - see Math for ML for a refresher

Chapter Map

#NoteTopicLevel
1LLM FundamentalsTokens, sampling parameters, context windows, model typesBeginner
2TokenizationBPE, WordPiece and Unigram, byte-level tokenizers, vocabulary size, the token tax on non-English text, numbers and glitch tokens, special tokensIntermediate
3Transformer ArchitectureEmbeddings, positional encoding, Pre-LN and RMSNorm, residuals, SwiGLUIntermediate
4Attention MechanismsQ/K/V math, multi-head, causal masking, FlashAttention, GQAIntermediate
5Model Architecture TypesEncoder, decoder, encoder-decoder, mixture of experts, model familiesIntermediate
6Modern ArchitecturesMLA, fine-grained MoE, local-global attention, attention sinks, SSM hybrids, native multimodalityAdvanced
7Failure Modes & Tricky IssuesForgetting, lost in the middle, hallucination, sycophancy, repetitionAdvanced
8Model LandscapeHow to read and choose among current models; dated snapshot of the major familiesBeginner
9Q&A Review Bank47 questions tagged Easy / Medium / HardAll levels

Where the Rest of the Old "LLM Models" Module Went

This module used to cover everything about LLMs. Its later chapters now live where they are taught in depth:

TopicNow in
Training & pretraining, GPU & hardwarePretraining at Scale
Fine-tuning, RLHF, DPO, GRPO, LoRAPost-Training & Reasoning
KV cache, inference optimization, production deploymentInference & Serving
Prompting mechanics (chat templates, constrained decoding)Prompt & Context Engineering

Review

  • Module quiz - every Check Yourself question in this module, in course order

Section Appendix

Summary & Key Terms - a quick recap of this section and its essential vocabulary.


Next Topic

Next: 02 - Python & Systems

⚡AI-assisted content - always verify, always explore multiple perspectives·