Quiz · 05 · Post-Training & Reasoning
17 questions from 4 pages
These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
SFT & Parameter-Efficient Fine-Tuning
Check yourself
0 / 5 answered
- LoRA with rank 8 is applied to a 4096 × 11008 projection. How many trainable parameters does it add?
- Why is B initialised to zero in LoRA?
- Your support bot must reflect a policy that changes weekly. Fine-tune or RAG?
- Why mask the prompt tokens out of the SFT loss?
- What does
prepare_model_for_kbit_trainingdo, and what happens if you skip it?
Preference Optimization
Check yourself
0 / 4 answered
- In the DPO loss, what does the reference model π_ref provide?
- You have 200,000 thumbs-up/thumbs-down ratings on single responses but no side-by-side pairs. Which method fits the data?
- Raising DPO's β from 0.05 to 0.5 mostly does what?
- Why can't DPO alone produce a strong reasoning model the way GRPO-style RL did?
RL for LLMs
Check yourself
0 / 4 answered
- What does GRPO use in place of PPO's value model?
- Every response in a GRPO group gets reward 1. What is the learning signal from that prompt?
- Why did RL with verifiable rewards make large-scale reasoning RL practical?
- Dr. GRPO removes GRPO's per-response length normalization. What bias did that normalization introduce?
Reasoning Models & Test-Time Compute
Check yourself
0 / 4 answered
- What rewards did DeepSeek-R1-Zero use?
- Why does the R1 pipeline include a small cold-start SFT stage before RL?
- You need a strong 7B reasoning model and have access to a much larger reasoning model. What did DeepSeek find works better: RL on the 7B model, or SFT on the large model's traces?
- Why shouldn't a model's visible chain of thought be treated as an audit trail?