Contents
Map

05 · Post-Training & Reasoning

RL for LLMs

View as:

Reinforcement Learning for LLMs

Online reinforcement learning - the model generates, gets scored, and updates on its own samples - is how assistants were first aligned (RLHF with PPO) and, since 2025, how reasoning models are trained (RL with verifiable rewards). This note covers the algorithms (PPO, GRPO and its fixes), the reward sources (reward models, verifiers, process rewards), reward hacking, and what RL training infrastructure looks like.

Learning objectives 70 min
By the end of this page you will be able to:
  • Explain the PPO-based RLHF loop, including the roles of the policy, reference, reward and value models
  • Derive GRPO's group-relative advantage and explain what it removes from PPO
  • Explain the fixes proposed by DAPO, Dr. GRPO and GSPO and the failure each targets
  • Compare outcome rewards, process rewards and verifiers, and recognize reward hacking
  • Describe the systems shape of RL training (rollout-dominated, generation engines, async pipelines)

The RL Loop for Language Models

flowchart LR
    P["📝 Prompts"] --> ROLL["🎲 Rollouts<br/>policy samples responses<br/>(inference engine)"]
    ROLL --> SCORE["🏅 Score<br/>reward model, verifier,<br/>or rules"]
    SCORE --> ADV["📐 Advantages<br/>value model (PPO) or<br/>group baseline (GRPO)"]
    ADV --> UPD["🔧 Policy update<br/>clipped objective + KL control"]
    UPD --> ROLL

    style ROLL fill:#d8dfe8,stroke:#b0bac8
    style SCORE fill:#e8e0d4,stroke:#c8b89a
    style ADV fill:#e8e2d9,stroke:#ccc4b8
    style UPD fill:#dde4dc,stroke:#b0c4b0

Language generation as RL: the state is the prompt plus tokens so far, an action is the next token, and the reward usually arrives once, at the end of the response. The policy gradient pushes up the log-probability of tokens in responses that scored better than expected - the "better than expected" part is the advantage, and how you estimate it is what separates the algorithms.


PPO-based RLHF

Concept

The InstructGPT-style pipeline keeps four models in play:

ModelRoleTrained?
PolicyGenerates responsesYes
Reference (frozen SFT)KL anchor - penalize drifting too farNo
Reward modelScores complete responsesNo (trained earlier)
Value model (critic)Predicts expected reward per token, for per-token advantages (GAE)Yes

PPO's clipped objective limits how far one update can move the policy:

ratio_t = π_θ(a_t|s_t) / π_old(a_t|s_t)
L = -E[ min( ratio_t · A_t,  clip(ratio_t, 1-ε, 1+ε) · A_t ) ]  +  β · KL(π_θ || π_ref)

PPO works but is heavy: two extra trainable-size models in memory, and the value model is hard to train when rewards come only at the end of long responses.


GRPO: Drop the Critic

Concept

Group Relative Policy Optimization (DeepSeekMath, 2024) replaces the learned value model with a statistical baseline. For each prompt, sample a group of G responses, score them, and use each one's standing within its group as the advantage:

A_i = ( r_i - mean(r_1..r_G) ) / std(r_1..r_G)        applied to every token of response i

The rest is PPO-style: a clipped ratio objective, plus a KL penalty to the reference model added to the loss. With binary correct/incorrect rewards this is intuitive - responses that beat the group average are reinforced, those below it are suppressed - and memory drops by a whole model.

RL with verifiable rewards (RLVR). For math (exact-match answers), code (unit tests), and format constraints, the reward is a program, not a learned model. That removes most reward-model hacking and made large-scale reasoning RL practical: DeepSeek-R1-Zero was trained with GRPO from a base model using only accuracy and format rewards (see Reasoning Models).

Fixes that followed

MethodProblem it targetsChange
DAPO (ByteDance Seed, 2025)Entropy collapse; wasted batches; long-response biasClip-higher (a larger upper clip bound so rare good tokens can grow); dynamic sampling (drop prompts where every sample got the same reward - zero advantage, zero signal); token-level loss averaging; soft penalties for overlong responses; no KL term
Dr. GRPO (2025)GRPO's normalizations bias learning - dividing by response length favours long wrong answers; dividing by group std overweights very easy/hard promptsRemove the per-response length normalization and the std normalization
GSPO (Qwen, 2025)Token-level importance ratios are noisy and destabilize large (especially MoE) modelsUse a sequence-level importance ratio and clip whole sequences

Reward Sources

RewardHow it worksStrengthsRisks
Learned outcome reward model (ORM)Scores a full response (Bradley-Terry)Works for open-ended quality (helpfulness, tone)Exploitable; rewards length and confident style
Verifier / rulesExact match, unit tests, compilers, format regexesHard to game; cheapOnly exists for checkable tasks; false positives in weak tests
Process reward model (PRM)Scores each reasoning stepDenser signal; can catch right-answer-wrong-reasoningStep labels are expensive (PRM800K) or noisy when automated; also hackable
LLM-as-judge with a rubricA strong model grades against explicit criteriaCovers tasks with no verifierJudge biases; the policy learns to please the judge

Reward hacking

The policy optimizes the reward you wrote, not the one you meant. Classic examples: responses padded with length because the reward model liked long answers; code that special-cases the unit tests; answers that print the expected format with no work. Defences: prefer verifiers, hold out an independent evaluation the policy never trains against, KL-regularize or cap training steps, inspect samples regularly, and make test suites adversarial.


What RL Training Looks Like as a System

  • Generation dominates. Most wall-clock time goes to sampling long responses, so RL frameworks embed a fast inference engine (vLLM or SGLang) next to the trainer and move weights between them every step.
  • Colocated vs disaggregated. Either the same GPUs alternate between generation and training, or separate pools run in parallel, with asynchronous (slightly off-policy) updates to keep both busy.
  • Long-tail rollouts. A few very long responses stall a synchronous batch; systems cap length, over-sample and drop stragglers, or run partial rollouts.
  • Frameworks: Hugging Face TRL (GRPOTrainer, PPOTrainer, DPOTrainer), verl (HybridFlow), OpenRLHF, and others. The Fine-Tuning Lab runs GRPO with TRL on a small model.

Check Yourself

Check yourself
0 / 4 answered
  1. What does GRPO use in place of PPO's value model?
  2. Every response in a GRPO group gets reward 1. What is the learning signal from that prompt?
  3. Why did RL with verifiable rewards make large-scale reasoning RL practical?
  4. Dr. GRPO removes GRPO's per-response length normalization. What bias did that normalization introduce?

Exercises

Exercise - Compute GRPO advantages

A prompt gets G = 6 responses with rewards [1, 0, 0, 1, 0, 0]. Compute each response's advantage (use population std). Then repeat for rewards [1, 1, 1, 1, 1, 0]. Which response in the second group gets the largest-magnitude advantage, and why is that desirable?

Solution

Group 1: mean = 1/3, std = 0.471 → correct responses +1.41, incorrect -0.71. Group 2: mean = 5/6, std = 0.373 → correct +0.45, the single incorrect -2.24. The lone failure on a mostly-solved prompt gets the strongest push away - the model learns most from its rare mistakes on problems it can usually solve.

Exercise - Spot the reward hack

Your code-RL run's reward (unit tests passing) rose from 40% to 85%, but a held-out benchmark improved only from 38% to 41%. List three hypotheses and how you would test each.

Hint

Read samples. Look for hard-coded expected outputs, disabled tests, or catching every exception.

Study Notes

Must-know:

  • RLHF-PPO needs policy, reference, reward and value models; clipped ratio objective + KL to reference
  • GRPO: sample G responses per prompt; advantage = (r - group mean) / group std; no critic
  • RLVR: rewards from verifiers/rules for math, code, format - the engine of 2025 reasoning models
  • DAPO: clip-higher, dynamic sampling, token-level loss, overlong shaping, no KL; Dr. GRPO: remove length and std normalization; GSPO: sequence-level ratios
  • Reward sources: ORM, verifier, PRM, rubric judge - all hackable to some degree; keep an independent held-out eval
  • RL systems are generation-bound: inference engine inside the loop, async/off-policy pipelines, long-tail rollouts

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·