Contents
Map

Quiz · 05 · Post-Training & Reasoning

17 questions from 4 pages

These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →

SFT & Parameter-Efficient Fine-Tuning

Check yourself
0 / 5 answered
  1. LoRA with rank 8 is applied to a 4096 × 11008 projection. How many trainable parameters does it add?
  2. Why is B initialised to zero in LoRA?
  3. Your support bot must reflect a policy that changes weekly. Fine-tune or RAG?
  4. Why mask the prompt tokens out of the SFT loss?
  5. What does prepare_model_for_kbit_training do, and what happens if you skip it?

Preference Optimization

Check yourself
0 / 4 answered
  1. In the DPO loss, what does the reference model π_ref provide?
  2. You have 200,000 thumbs-up/thumbs-down ratings on single responses but no side-by-side pairs. Which method fits the data?
  3. Raising DPO's β from 0.05 to 0.5 mostly does what?
  4. Why can't DPO alone produce a strong reasoning model the way GRPO-style RL did?

RL for LLMs

Check yourself
0 / 4 answered
  1. What does GRPO use in place of PPO's value model?
  2. Every response in a GRPO group gets reward 1. What is the learning signal from that prompt?
  3. Why did RL with verifiable rewards make large-scale reasoning RL practical?
  4. Dr. GRPO removes GRPO's per-response length normalization. What bias did that normalization introduce?

Reasoning Models & Test-Time Compute

Check yourself
0 / 4 answered
  1. What rewards did DeepSeek-R1-Zero use?
  2. Why does the R1 pipeline include a small cold-start SFT stage before RL?
  3. You need a strong 7B reasoning model and have access to a much larger reasoning model. What did DeepSeek find works better: RL on the 7B model, or SFT on the large model's traces?
  4. Why shouldn't a model's visible chain of thought be treated as an audit trail?