Code Lab 02 - SFT, DPO and GRPO with TRL
Run the three stages of modern post-training on a small open model, with LoRA adapters and Hugging Face TRL: supervised fine-tuning on chat demonstrations, DPO on preference pairs, then GRPO with verifiable math rewards. One script, one command per stage.
← Back to Overview: Fine-Tuning Lab · Concepts: Preference Optimization · RL for LLMs
- Run SFT, DPO and GRPO with TRL's trainers and explain what each stage's dataset looks like
- Read DPO training metrics (loss starting at ln 2, reward margins and accuracies) and GRPO metrics (per-reward means, group std)
- Write verifiable reward functions and debug them before spending GPU time
- Decide when to merge LoRA adapters between stages
What's In This Lab
| Property | Detail |
|---|---|
| Base model | Qwen/Qwen3-0.6B by default (any causal LM with a chat template works) |
| Stages | sft on trl-lib/Capybara → dpo on trl-lib/ultrafeedback_binarized → grpo on GSM8K |
| Adapters | LoRA on all linear layers; merged after each stage so the next stage (or vLLM) loads a plain model |
| GRPO rewards | correctness_reward (final #### <number> matches the reference: 1.0) + format_reward (answer line in the required format: 0.2) |
| Verified | All three stages run end to end on a laptop CPU with --smoke (tiny random model, 8 examples, 2 steps) using TRL 1.14 / transformers 5.17 / peft 0.21. Full runs are not GPU-verified here - expect roughly an hour per stage on one 24 GB GPU at the default sizes. |
| Files | 02-SFT-DPO-GRPO-with-TRL/{post_train.py, requirements.txt} |
flowchart LR
B["🧱 Qwen3-0.6B"] --> S["📝 SFT<br/>Capybara chats<br/>LoRA → merge"]
S --> D["⚖️ DPO<br/>UltraFeedback pairs<br/>β = 0.1 → merge"]
D --> G["🎯 GRPO<br/>GSM8K, G = 8 samples / prompt<br/>correctness + format rewards → merge"]
G --> E["🧪 Evaluate<br/>GSM8K test accuracy,<br/>chat quality spot-checks"]
style S fill:#d8dfe8,stroke:#b0bac8
style D fill:#dde4dc,stroke:#b0c4b0
style G fill:#e8e0d4,stroke:#c8b89a
style E fill:#ddd8e4,stroke:#b8b0c8
Run It
cd 06-Fine-Tuning-Lab/CodeLabs/02-SFT-DPO-GRPO-with-TRL
pip install -r requirements.txt
# Smoke-test every stage first (minutes on a laptop CPU)
for s in sft dpo grpo; do python post_train.py $s --model trl-internal-testing/tiny-Qwen3ForCausalLM --smoke; done
# Real runs (GPU)
python post_train.py sft --output-dir runs/sft
python post_train.py dpo --model runs/sft/merged --output-dir runs/dpo
python post_train.py grpo --model runs/dpo/merged --output-dir runs/grpo
Walkthrough - What to Look At
- Dataset shapes. SFT rows are
messageslists; DPO rows areprompt/chosen/rejected; GRPO rows are aprompt(system + user messages) plus asolutioncolumn that TRL passes to the reward functions as a keyword argument. - DPO's first loss is ln 2 ≈ 0.693. At step 0 the policy equals the reference, so the implicit reward margin is zero and
-log σ(0) = ln 2. Watchrewards/marginsgrow andrewards/accuraciesrise above 0.5. - No second model for DPO's reference. With a LoRA adapter, TRL computes reference log-probabilities by disabling the adapter - the frozen base is the reference.
- GRPO's group.
num_generationsis G: completions sampled per prompt to form the group baseline. Prompts where all G samples get the same reward contribute nothing - watch the per-rewardstdin the logs. - TRL's GRPO defaults are DAPO-style.
GRPOConfigdefaults to a token-level ("dapo") loss andbeta = 0.0- no KL penalty. Setbeta > 0to add one. - Test rewards before training. The reward functions are plain Python - call them on hand-written completions (see the exercise) before spending GPU hours.
Check Yourself
- Your DPO run logs a loss of 0.693 at step 0. What does that tell you?
- In GRPO logs, the correctness reward's std within groups is 0 for most prompts. What is happening, and what should you change?
- Why merge the LoRA adapter after each stage?
Exercises
Import correctness_reward and format_reward from post_train.py and write four test completions: correct and well-formatted; correct with a comma in the number (#### 1,200); wrong; correct answer followed by extra text. Predict each reward, then check. Is rewarding format separately a risk (can the model learn to output #### 0 for the format bonus)?
Solution
Expected: (1.0, 0.2), (1.0, 0.2), (0.0, 0.0 or 0.2 depending on format), (1.0, 0.0). A format-only strategy earns 0.2 per sample, far below the 1.0 for correct answers, so it is not optimal once the model can solve some problems - but keep the format reward small relative to correctness for exactly this reason.
Evaluate the base model, and each stage's merged output, on 200 GSM8K test questions using the same prompt and answer extraction. Plot accuracy by stage and describe which stage moved it most. Then spot-check chat quality on 10 open-ended prompts - did GRPO on math hurt general chat?
Rerun GRPO with beta=0.04. Compare reward curves and GSM8K accuracy with the default (beta=0.0), and the model's behaviour on non-math prompts.
References
- Hugging Face, TRL documentation -
SFTTrainer,DPOTrainer,GRPOTrainer - Rafailov et al., Direct Preference Optimization (2023)
- Shao et al., DeepSeekMath (GRPO) (2024); Yu et al., DAPO (2025)
- Cobbe et al., Training Verifiers to Solve Math Word Problems (GSM8K) (2021)
- Cui et al., UltraFeedback (2023)
Last reviewed: 2026-09