Post-Training & Reasoning - Q&A Review Bank
21 Q&A pairs. Tags:
[Easy]= conceptual recall,[Medium]= design decisions and trade-offs,[Hard]= derivations, quantitative reasoning or system design.
- Answer each question from memory before revealing the answer, across: Pipeline and SFT; Preference Optimization; Reinforcement Learning; Reasoning Models; Preference Tuning in Practice
- Explain the reasoning behind each answer - the mechanism or trade-off - not only the fact
- Identify the chapters you are weakest on and revisit them before the module quiz
- The concept notes of this module
Pipeline and SFT
Q1 [Easy] What are the typical stages of post-training for a chat or reasoning model?
Supervised fine-tuning on demonstrations, then preference tuning (DPO-family on chosen/rejected pairs) and/or online RL (PPO or GRPO with reward models or verifiers). Reasoning models add long-chain-of-thought RL with verifiable rewards, often wrapped in rejection-sampling SFT rounds.
Q2 [Medium] Why is the loss masked to the response tokens during SFT?
The model should learn to produce responses, not to predict the user's prompt or the system prompt. Masking prompt tokens (label -100) focuses the gradient on the assistant turns and avoids teaching the model to imitate user text.
Preference Optimization
Q3 [Medium] Write the Bradley-Terry reward-model loss and explain it.
L = -log σ(r(x, y_w) - r(x, y_l)). It maximizes the modelled probability that the chosen response beats the rejected one, which depends only on the difference of their scalar rewards.
Q4 [Hard] Sketch how DPO removes the reward model.
The KL-regularized RLHF objective has the closed-form optimum π*(y|x) ∝ π_ref(y|x)·exp(r(x,y)/β). Rearranging gives r(x,y) = β·log(π(y|x)/π_ref(y|x)) + const. Substituting that implicit reward into the Bradley-Terry loss cancels the constant and yields a loss purely in terms of the policy and reference log-probabilities: -log σ(β[log-ratio(chosen) - log-ratio(rejected)]).
Q5 [Medium] What does β control in DPO?
The strength of the implicit KL constraint: larger β keeps the policy closer to the reference model; smaller β allows larger moves (more gain, more risk of degradation or reward over-optimization).
Q6 [Medium] When would you choose KTO, SimPO or ORPO over DPO?
KTO when you only have unpaired good/bad labels (e.g. thumbs up/down). SimPO when you want to drop the reference model (memory) and counter length exploitation with length-normalized rewards and a margin. ORPO when you want to fold SFT and preference tuning into one reference-free stage.
Q7 [Medium] What are DPO's main limitations?
It is offline, so it cannot discover behaviour beyond its fixed pairs; it degrades when pairs come from a distribution far from the policy's; it can reduce the likelihood of both chosen and rejected responses; and it tends to lengthen outputs. Iterative/online DPO (regenerating pairs from the current policy) mitigates the first two.
Reinforcement Learning
Q8 [Medium] Which models does PPO-based RLHF keep in memory, and why?
The policy (trained), a frozen reference (KL anchor), a reward model (scores responses) and a value model (per-token value estimates for advantages). That is roughly four model copies - the main reason GRPO's removal of the value model matters.
Q9 [Medium] How does GRPO compute advantages?
Sample G responses per prompt, score them, and set each response's advantage to (reward - group mean) / group std, applied to all its tokens. The group acts as the baseline, so no value model is needed.
Q10 [Hard] A GRPO group of 4 gets rewards [1, 1, 0, 0]. What are the advantages, and what if all four were 1?
Mean 0.5, std 0.5 → +1 for the correct responses, -1 for the incorrect. If all are 1, the std is zero and every advantage is zero (implementations add ε), so the prompt contributes no gradient - which is why DAPO's dynamic sampling filters such prompts out.
Q11 [Medium] What is RLVR, and why did it matter?
RL with verifiable rewards: rewards come from programs - exact-match answers, unit tests, format checks - rather than learned reward models. It is hard to exploit and cheap, which made long reasoning-focused RL runs practical (the Tülu 3 paper named it; DeepSeek-R1 used it at scale).
Q12 [Hard] Name three changes DAPO made to GRPO and the problem each solves.
Clip-higher (a larger upper clip bound) prevents entropy collapse by letting low-probability good tokens grow; dynamic sampling drops prompts whose samples all share one reward, since they carry zero advantage; token-level loss averaging stops long responses being under-weighted; plus overlong reward shaping and removing the KL term.
Q13 [Medium] What is reward hacking? Give two examples and two defences.
The policy maximizes the reward as written rather than as intended - e.g. padding answers because the reward model likes length, or special-casing unit tests instead of fixing code. Defences: prefer robust verifiers and adversarial tests, and keep an independent held-out evaluation; also KL regularization, step caps and regular sample inspection.
Q14 [Medium] Outcome reward model vs process reward model?
An ORM scores the complete response; a PRM scores each reasoning step, giving denser credit assignment and catching right-answer-wrong-reasoning. PRMs need expensive step labels (or noisy automated ones) and can also be hacked; DeepSeek-R1 reported they did not pay off relative to simple outcome rewards in their setting.
Q15 [Hard] Why is RL training dominated by generation, and how do systems cope?
Each step needs many long sampled responses, and autoregressive decoding is slow relative to a training step. Frameworks embed a fast inference engine (vLLM/SGLang) in the loop, sync weights to it every step, run generation and training on separate or time-shared GPU pools, use asynchronous slightly off-policy updates, and cap or drop long-tail rollouts.
Reasoning Models
Q16 [Medium] Describe the DeepSeek-R1 training pipeline.
Cold-start SFT on thousands of curated long-CoT examples; reasoning-focused GRPO with accuracy, format and language-consistency rewards; rejection sampling to build ~600K reasoning + ~200K general SFT samples and a second SFT; then RL across all scenarios combining verifiable rewards and reward models for helpfulness and harmlessness.
Q17 [Medium] What happened in DeepSeek-R1-Zero, and why wasn't it shipped as the main model?
Pure GRPO from the base model with only accuracy and format rewards produced growing response lengths and emergent self-verification and backtracking, with strong reasoning scores - but outputs had poor readability and mixed languages, which the cold-start SFT stage in R1 fixed.
Q18 [Medium] Why distil reasoning into small models instead of running RL on them?
DeepSeek found SFT on a strong reasoning model's traces gave small models better reasoning than large-scale RL applied to the small models directly - RL is expensive and small models explore poorly, while distillation transfers behaviour the large model already discovered.
Q19 [Medium] List four ways to spend compute at test time.
Longer thinking (more reasoning tokens / higher effort), self-consistency (majority vote over samples), best-of-N with a verifier or reward model, and sequential self-revision; plus search over reasoning steps with a step-level scorer.
Q20 [Medium] What are the main failure modes of reasoning models in production?
Overthinking easy inputs (cost and latency), chains of thought that don't faithfully reflect the real reason for an answer, reward-hacking behaviours learned during RL, and smaller gains on open-ended tasks without verifiable rewards. Effort controls and per-route evaluation are the practical levers.
Preference Tuning in Practice
Q21 [Medium] Why did DPO become a common default over PPO-based RLHF, and where does online RL still win?
DPO reformulates preference optimization as a supervised loss directly on (chosen, rejected) pairs - no separate reward model training, no PPO. RLHF/PPO challenges: reward model is a proxy (can be gamed), PPO is unstable (KL constraint must be tuned carefully), requires two separate training stages. DPO: one training stage, same stability as SFT, competitive results on chat-preference benchmarks. Llama 3's post-training, for example, used SFT + rejection sampling + DPO rather than PPO. The limit: DPO is offline (it learns from a fixed set of pairs), so it cannot discover behaviours beyond that data. Reasoning models brought online RL back - GRPO-style RL with verifiable rewards (DeepSeek-R1, 2025) trains on the model's own fresh samples scored by a checker, which is how long chain-of-thought reasoning is elicited.