Contents
Map

Quiz · 06 · Fine-Tuning Lab

30 questions from 6 pages

These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →

HuggingFace Ecosystem

Check yourself
0 / 5 answered
  1. Which padding side does batched generation with a decoder-only model need?
  2. You set pad_token = eos_token and build labels by setting -100 wherever input_ids == pad_token_id. What goes wrong?
  3. Why does a mismatched tokenizer/model pairing not throw an error?
  4. What happens if you skip apply_chat_template and just concatenate strings yourself for an instruction model?
  5. Why is datasets Arrow-backed instead of just loading a Python list/dict?

LoRA & QLoRA Hands-On

Check yourself
0 / 6 answered
  1. lora_alpha = 32 and r = 16. What multiplies the adapter's output?
  2. You need one base model to serve 40 customer-specific fine-tunes. What do you do?
  3. Why does prepare_model_for_kbit_training matter for QLoRA?
  4. When would you choose adapter-only serving over merging?
  5. What's the tradeoff of a higher LoRA rank?
  6. Can you merge a QLoRA adapter straight back into the 4-bit quantized base?

Instruction Data & Training Runs

Check yourself
0 / 7 answered
  1. Why does the training loop use DataCollatorForSeq2Seq rather than default_data_collator?
  2. Training loss falls to near zero during epoch 1 on a 500-example dataset. Most likely?
  3. What does setting a label to -100 actually do mechanically?
  4. Why mask the prompt tokens instead of just weighting them lower?
  5. You see loss oscillating without trending down - what are the two most likely first things to try?
  6. Why is overfitting harder to spot in fine-tuning than in typical ML training?
  7. Training loss looks great (steadily decreasing, ends near zero) but the fine-tuned model performs worse than the base model on held-out prompts. What happened?

Benchmarking Base vs Tuned

Check yourself
0 / 6 answered
  1. Why call torch.cuda.synchronize() before reading the timer?
  2. A tuned model beats few-shot prompting by 1.5 points on 60 eval items. What do you conclude?
  3. Why exclude the first generate() call from a latency benchmark?
  4. Why measure peak VRAM instead of computing it from parameter count?
  5. A fine-tune scores well on the eval set but users report worse real-world behavior - what's the most likely explanation?
  6. Give one clear signal that RAG, not fine-tuning, is the right fix.

QLoRA Fine-Tune

Check yourself
0 / 3 answered
  1. The benchmark compares base and tuned models. Where must the evaluation examples come from?
  2. Why does benchmark.py run warm-up generations before timing?
  3. The tuned model is better on the task but 20% slower. What is the most likely reason if the adapter was not merged?

SFT, DPO & GRPO with TRL

Check yourself
0 / 3 answered
  1. Your DPO run logs a loss of 0.693 at step 0. What does that tell you?
  2. In GRPO logs, the correctness reward's std within groups is 0 for most prompts. What is happening, and what should you change?
  3. Why merge the LoRA adapter after each stage?