Quiz · 06 · Fine-Tuning Lab
30 questions from 6 pages
These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
HuggingFace Ecosystem
Check yourself
0 / 5 answered
- Which padding side does batched generation with a decoder-only model need?
- You set pad_token = eos_token and build labels by setting -100 wherever input_ids == pad_token_id. What goes wrong?
- Why does a mismatched tokenizer/model pairing not throw an error?
- What happens if you skip
apply_chat_templateand just concatenate strings yourself for an instruction model? - Why is
datasetsArrow-backed instead of just loading a Python list/dict?
LoRA & QLoRA Hands-On
Check yourself
0 / 6 answered
- lora_alpha = 32 and r = 16. What multiplies the adapter's output?
- You need one base model to serve 40 customer-specific fine-tunes. What do you do?
- Why does
prepare_model_for_kbit_trainingmatter for QLoRA? - When would you choose adapter-only serving over merging?
- What's the tradeoff of a higher LoRA rank?
- Can you merge a QLoRA adapter straight back into the 4-bit quantized base?
Instruction Data & Training Runs
Check yourself
0 / 7 answered
- Why does the training loop use DataCollatorForSeq2Seq rather than default_data_collator?
- Training loss falls to near zero during epoch 1 on a 500-example dataset. Most likely?
- What does setting a label to
-100actually do mechanically? - Why mask the prompt tokens instead of just weighting them lower?
- You see loss oscillating without trending down - what are the two most likely first things to try?
- Why is overfitting harder to spot in fine-tuning than in typical ML training?
- Training loss looks great (steadily decreasing, ends near zero) but the fine-tuned model performs worse than the base model on held-out prompts. What happened?
Benchmarking Base vs Tuned
Check yourself
0 / 6 answered
- Why call torch.cuda.synchronize() before reading the timer?
- A tuned model beats few-shot prompting by 1.5 points on 60 eval items. What do you conclude?
- Why exclude the first
generate()call from a latency benchmark? - Why measure peak VRAM instead of computing it from parameter count?
- A fine-tune scores well on the eval set but users report worse real-world behavior - what's the most likely explanation?
- Give one clear signal that RAG, not fine-tuning, is the right fix.
QLoRA Fine-Tune
Check yourself
0 / 3 answered
- The benchmark compares base and tuned models. Where must the evaluation examples come from?
- Why does benchmark.py run warm-up generations before timing?
- The tuned model is better on the task but 20% slower. What is the most likely reason if the adapter was not merged?
SFT, DPO & GRPO with TRL
Check yourself
0 / 3 answered
- Your DPO run logs a loss of 0.693 at step 0. What does that tell you?
- In GRPO logs, the correctness reward's std within groups is 0 for most prompts. What is happening, and what should you change?
- Why merge the LoRA adapter after each stage?