Contents
Map

06 · Fine-Tuning Lab

QLoRA Fine-Tune

View as:

Code Lab 01 - QLoRA Fine-Tune

A real, runnable QLoRA fine-tune of a small open instruction model on a compact instruction dataset - loaded in 4-bit via bitsandbytes, adapted with a peft LoRA config, trained with prompt/completion masking, and benchmarked against the un-tuned base model on quality, latency, and VRAM.

← Back to Overview: Fine-Tuning Lab · Back to Concepts: LoRA & QLoRA Hands-On · Instruction Data & Training Runs · Benchmarking Base vs Tuned

Learning objectives 2-3 hours (plus GPU time)
By the end of this page you will be able to:
  • Run a QLoRA fine-tune end to end - 4-bit base, LoRA adapters, masked prompt/completion loss, held-out split
  • Benchmark the tuned model against its base on quality, tokens per second and peak VRAM
  • Decide from the benchmark whether the fine-tune is worth deploying
Prerequisites

What's In This Lab

PropertyDetail
TaskInstruction-following fine-tune (support-ticket-style responses; swappable for any instruction/response dataset)
Base modelA small open ~1-2B instruction model (default: Qwen/Qwen3-1.7B), reachable on a single consumer/free-tier GPU
Quantization4-bit NF4 base weights via bitsandbytes, BF16 LoRA adapters
Adapterpeft LoraConfig with target_modules="all-linear" (r = 16, alpha = 32)
TrainingMasked prompt/completion loss, transformers.Trainer, checkpointed adapter output
BenchmarkBase vs adapter-merged model: quality proxy, tokens/sec, peak VRAM
ComplexityIntermediate-Advanced
Files01-QLoRA-Fine-Tune/{train_qlora.py, benchmark.py, requirements.txt, README.mdx}

Architecture

flowchart TD
    subgraph Data["Data Pipeline"]
        direction LR
        Raw["Instruction dataset\n(JSONL)"] --> Split["train/eval split\n(held out)"]
        Split --> Fmt["Format + tokenize\n+ mask prompt tokens"]
    end
    subgraph Load["Model Loading"]
        direction LR
        HF["AutoModelForCausalLM\n.from_pretrained"] --> BNB["BitsAndBytesConfig\nNF4 + double quant"]
        BNB --> Prep["prepare_model_for_kbit_training"]
    end
    subgraph Train["Training"]
        direction LR
        Lora["LoraConfig\nr=16, alpha=32"] --> Peft["get_peft_model"]
        Peft --> Loop["Trainer.train()\nmasked loss"]
    end
    subgraph Bench["Benchmark"]
        direction LR
        Merge["merge_and_unload()"] --> Compare["Base vs Tuned\nquality / tok-s / VRAM"]
        Compare --> Table["Comparison table"]
    end

    Fmt --> Loop
    Prep --> Peft
    Loop -->|save adapter| Merge

    style Raw fill:#d8dfe8,stroke:#b0bac8
    style Split fill:#d8dfe8,stroke:#b0bac8
    style Fmt fill:#d8dfe8,stroke:#b0bac8
    style HF fill:#e8e0d4,stroke:#c8b89a
    style BNB fill:#e8e0d4,stroke:#c8b89a
    style Prep fill:#e8e0d4,stroke:#c8b89a
    style Lora fill:#dde4dc,stroke:#b0c4b0
    style Peft fill:#dde4dc,stroke:#b0c4b0
    style Loop fill:#dde4dc,stroke:#b0c4b0
    style Merge fill:#ddd8e4,stroke:#b8b0c8
    style Compare fill:#ddd8e4,stroke:#b8b0c8
    style Table fill:#ddd8e4,stroke:#b8b0c8

Running

cd 06-Fine-Tuning-Lab/CodeLabs/01-QLoRA-Fine-Tune
pip install -r requirements.txt
python train_qlora.py --epochs 3 --output-dir ./qlora-adapter
python benchmark.py --adapter-dir ./qlora-adapter

Full details, code walkthrough, and what each part demonstrates: see README.

Check Yourself

Check yourself
0 / 3 answered
  1. The benchmark compares base and tuned models. Where must the evaluation examples come from?
  2. Why does benchmark.py run warm-up generations before timing?
  3. The tuned model is better on the task but 20% slower. What is the most likely reason if the adapter was not merged?

Exercises

Exercise - Merge and re-benchmark

Load the base in bf16, attach the trained adapter, merge it with merge_and_unload(), save the merged model and re-run benchmark.py. Compare quality and tokens per second with the unmerged 4-bit setup.

Solution

Quality should match within noise (the merge is exact in bf16); tokens per second usually improves because the merged model has no adapter path or 4-bit dequantisation overhead - at the cost of a larger artifact and higher memory.

Exercise - Is it worth it?

Add a third arm to the benchmark: the base model with a 3-shot prompt. Compare all three on quality (with confidence intervals), input tokens per request and cost per 1,000 requests, then write a one-paragraph recommendation.

Solution

Few-shot usually closes part of the gap at the cost of longer prompts. If the tuned model's quality lead over few-shot is within the confidence interval, the recommendation is to ship the prompt; if it is clearly larger and volume is high, the fine-tune pays back (see the break-even exercise in Benchmarking Base vs Tuned).

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·