06 - Fine-Tuning & PEFT Lab
Hands-on fine-tuning with the Hugging Face stack: loading models and chat templates, LoRA and QLoRA configuration, instruction data with masked losses, reading loss curves, and benchmarking a fine-tune against its base - then the full SFT, DPO and GRPO sequence with TRL.
Learning objectives 8-10 hours (notes + labs, plus GPU time)
By the end of this module you will be able to:- Load models, tokenizers and chat templates with transformers, and prepare datasets with the datasets library
- Configure and run LoRA and QLoRA fine-tunes, and merge or serve adapters
- Format instruction data with masked losses and diagnose training runs from loss curves
- Benchmark a fine-tune against base and prompted baselines on quality, latency, VRAM and cost, and decide whether it is worth deploying
- Run SFT, DPO and GRPO with TRL
Prerequisites
Chapter Map
| # | File | Topic | Difficulty |
|---|---|---|---|
| 1 | HuggingFace Ecosystem | AutoModelForCausalLM/AutoTokenizer, tokenization, chat templates, datasets, the Hub | Beginner |
| 2 | LoRA & QLoRA Hands-On | LoRA config (rank/alpha/dropout/target modules), QLoRA (NF4, double quant), merging vs serving adapters | Intermediate |
| 3 | Instruction Data & Training Runs | Formatting instruction data, prompt/completion masking, reading loss curves, a hand-written PEFT training loop | Intermediate |
| 4 | Benchmarking Base vs Tuned | Held-out eval sets, quality delta, measured latency/tokens-per-sec, VRAM and cost-per-1k-requests | Advanced |
| 5 | Q&A Review Bank | 27 Q&A pairs across all topics in this module | All levels |
Recommended Learning Paths
Path A: First Fine-Tune, Start to Finish
- HuggingFace Ecosystem - learn the tools you'll be using
- LoRA & QLoRA Hands-On - understand what you're configuring and why
- Instruction Data & Training Runs - format your data and run the training loop
- QLoRA Fine-Tune Code Lab - do it end-to-end on a real small model
- SFT, DPO & GRPO with TRL - run the full post-training sequence
- Benchmarking Base vs Tuned - prove the fine-tune actually helped
Path B: Interview Preparation (Accelerated)
- LoRA & QLoRA Hands-On - config parameters and merge-vs-serve questions are very common
- Instruction Data & Training Runs - loss masking and loss-curve reading come up constantly
- Benchmarking Base vs Tuned - "how do you know it worked" is a favorite follow-up question
- Q&A Review Bank - drill all 27 questions
Path C: Production Engineering (Advanced)
- LoRA & QLoRA Hands-On - merging adapters vs adapter-swapping at serve time
- Benchmarking Base vs Tuned - cost-per-1k-requests and when fine-tuning is the wrong call
- QLoRA Fine-Tune Code Lab -
train_qlora.py+benchmark.pyas a reusable template
Resources
- Q&A Review Bank - 27 Q&A pairs in this module
- Module quiz - every Check Yourself question in this module
- QLoRA Fine-Tune Code Lab - hands-on PEFT + QLoRA fine-tune with a base-vs-tuned benchmark, formatting and loss masking done by hand
- SFT, DPO & GRPO with TRL Code Lab - the three post-training stages with TRL, including verifiable math rewards
Key Cross-References
- LoRA math (ΔW = BA, rank/alpha derivation) and the map of preference tuning and RL → Post-Training: SFT & PEFT - this module builds hands-on tooling on top of those foundations, it does not re-derive the math
- VRAM estimation and precision formats → Pretraining: GPU Memory & Hardware
- The raw-PyTorch training loop this module's hand-written PEFT loop extends → PyTorch Fundamentals: Training Loop From Scratch
- Checkpointing and mixed precision used during training runs → PyTorch Fundamentals: Checkpointing & Mixed Precision
Section Appendix
Summary & Key Terms - a quick recap of this section and its essential vocabulary.
Next Topic
Previous: 05 - Post-Training & Reasoning · Next: 07 - Evaluation & Benchmarks