05 - Post-Training & Reasoning
A pretrained base model continues text; it does not follow instructions, refuse harmful requests, or reason step by step on demand. Post-training adds those behaviours: supervised fine-tuning on demonstrations, preference optimization, and reinforcement learning - including the RL-with-verifiable-rewards recipe behind today's reasoning models.
Learning objectives 6-8 hours
By the end of this module you will be able to:- Explain the post-training pipeline from SFT to preference tuning to RL, and what each stage changes in the model
- Compare RLHF with PPO, DPO-family methods and GRPO by what they need (reward model, critic, online samples) and what they cost
- Explain how LoRA and QLoRA make fine-tuning affordable, and what they do not fix
- Describe how RL with verifiable rewards produces long chain-of-thought reasoning
Prerequisites
- LLM Foundations
- Pretraining at Scale - especially the training memory math
Chapter Map
| # | Note | Topic | Level |
|---|---|---|---|
| 1 | SFT & Parameter-Efficient Fine-Tuning | When to fine-tune, SFT and loss masking, LoRA and QLoRA, PEFT comparison, evaluating fine-tunes; a map of preference tuning and RL | Intermediate |
| 2 | Preference Optimization | Bradley-Terry reward models, DPO derivation and β, IPO, KTO, SimPO, ORPO, offline vs online | Advanced |
| 3 | RL for LLMs | PPO-RLHF, GRPO, RLVR, DAPO / Dr. GRPO / GSPO, reward sources and hacking, RL systems | Advanced |
| 4 | Reasoning Models & Test-Time Compute | R1-Zero and R1 recipes, distillation, self-consistency, best-of-N, budget forcing, effort dials, failure modes | Advanced |
| 5 | Q&A Review Bank | 21 questions across SFT, preference tuning, RL and reasoning | All levels |
Hands-on practice for this module is in the Fine-Tuning Lab.
Review
- Q&A Review Bank - 21 questions for this module
Section Appendix
Summary & Key Terms - a quick recap of this section and its essential vocabulary.
Next Topic
Previous: 04 - Pretraining at Scale · Next: 06 - Fine-Tuning Lab