Contents
Map

05 · Post-Training & Reasoning

Overview

View as:

05 - Post-Training & Reasoning

A pretrained base model continues text; it does not follow instructions, refuse harmful requests, or reason step by step on demand. Post-training adds those behaviours: supervised fine-tuning on demonstrations, preference optimization, and reinforcement learning - including the RL-with-verifiable-rewards recipe behind today's reasoning models.

Learning objectives 6-8 hours
By the end of this module you will be able to:
  • Explain the post-training pipeline from SFT to preference tuning to RL, and what each stage changes in the model
  • Compare RLHF with PPO, DPO-family methods and GRPO by what they need (reward model, critic, online samples) and what they cost
  • Explain how LoRA and QLoRA make fine-tuning affordable, and what they do not fix
  • Describe how RL with verifiable rewards produces long chain-of-thought reasoning
Prerequisites

Chapter Map

#NoteTopicLevel
1SFT & Parameter-Efficient Fine-TuningWhen to fine-tune, SFT and loss masking, LoRA and QLoRA, PEFT comparison, evaluating fine-tunes; a map of preference tuning and RLIntermediate
2Preference OptimizationBradley-Terry reward models, DPO derivation and β, IPO, KTO, SimPO, ORPO, offline vs onlineAdvanced
3RL for LLMsPPO-RLHF, GRPO, RLVR, DAPO / Dr. GRPO / GSPO, reward sources and hacking, RL systemsAdvanced
4Reasoning Models & Test-Time ComputeR1-Zero and R1 recipes, distillation, self-consistency, best-of-N, budget forcing, effort dials, failure modesAdvanced
5Q&A Review Bank21 questions across SFT, preference tuning, RL and reasoningAll levels

Hands-on practice for this module is in the Fine-Tuning Lab.

Review

Section Appendix

Summary & Key Terms - a quick recap of this section and its essential vocabulary.


Next Topic

Previous: 04 - Pretraining at Scale · Next: 06 - Fine-Tuning Lab

⚡AI-assisted content - always verify, always explore multiple perspectives·