Contents
Map

05 · Post-Training & Reasoning

Appendix - Summary & Key Terms

View as:

Appendix - Post-Training & Reasoning

What We Learned

  • SFT teaches desired responses; PEFT updates a small fraction of parameters.
  • Preference optimization learns from comparisons; online RL learns from sampled outcomes and rewards.
  • Reference-policy constraints limit drift but do not guarantee safety or correctness.
  • Reasoning models spend compute on intermediate work; more thinking is not always better.
  • Evaluate task gains, forgetting and reward hacking on held-out data.

Key Acronyms, Concepts & Jargon

TermShort meaning
SFTSupervised Fine-Tuning: training on examples of desired responses.
PEFTParameter-Efficient Fine-Tuning: adapting a model with few trainable parameters.
LoRA / QLoRALow-Rank Adaptation / Quantized LoRA: train low-rank updates / do so on a quantized base.
RL / RLHFReinforcement Learning / RL from Human Feedback: learn from rewards / human-derived preferences.
Reward modelPredicts preference or quality to supply a training signal.
DPODirect Preference Optimization: learns from preferred/rejected pairs without an explicit RL rollout loop.
PPOProximal Policy Optimization: constrains policy updates during reward optimization.
GRPOGroup Relative Policy Optimization: uses group-relative rewards without a separate value critic.
RLVRReinforcement Learning with Verifiable Rewards: trains against checkable outcomes.
KL penaltyDiscourages a policy from drifting away from a reference distribution.
CoT / test-time computeChain of Thought / inference work spent reasoning, searching or checking.
DistillationTrains a smaller or cheaper model using a stronger model's outputs.
Reward hacking / forgettingExploiting the reward without solving the task / losing earlier capabilities.

Back to section overview

⚡AI-assisted content - always verify, always explore multiple perspectives·