Quantized Inference
The One-Line Definition
Inference-time quantization (GPTQ, AWQ) and training-time quantization (NF4/bitsandbytes) both shrink model weights to fewer bits, but they optimize for different things - GPTQ/AWQ are calibrated once, offline, purely to make a frozen model cheaper and faster to serve, while NF4/bitsandbytes is designed to make a model cheap enough to train against, with speed as a secondary concern.
- Distinguish training-time quantization (NF4) from serving formats (FP8, INT8, GPTQ, AWQ, FP4)
- Choose a serving precision for a GPU generation and task sensitivity
- Produce and serve a quantized checkpoint with llm-compressor and vLLM, and gate it on a quality evaluation
If the Fine-Tuning Lab (module 06) taught you to quantize a model so you could afford to fine-tune it on a single GPU, this page is about a related but distinct decision: quantizing an already-trained model purely to make it cheaper and faster to run for everyone who calls it in production. The tools and priorities are different even though the underlying idea - "use fewer bits per weight" - sounds the same.
This page assumes the quantization fundamentals (NF4 bin placement, VRAM formulas, INT8/INT4 tradeoffs) from Pretraining: GPU & Hardware and the QLoRA training-time setup from Fine-Tuning Lab: LoRA & QLoRA Hands-On. It is specifically about serving-side quantization - FP8 and FP4 on current GPUs, and INT4 weight-only formats (GPTQ, AWQ) - and when to reach for each.
flowchart LR
FP["๐ฏ Full-precision\nfine-tuned model"]
Cal["๐ Calibration set\n(few hundred samples)"]
GA["๐ข llm-compressor\nFP8 / AWQ / GPTQ / NVFP4\nquantize offline"]
Art["๐ฆ Quantized checkpoint\n(FP8 โ 50%, INT4/FP4 โ 25-30%\nof BF16 size)"]
Serve["๐ vLLM / SGLang serve\nwith quantized kernels"]
FP --> Cal --> GA --> Art --> Serve
style FP fill:#d8dfe8,stroke:#b0bac8
style Cal fill:#e8e0d4,stroke:#c8b89a
style GA fill:#dde4dc,stroke:#b0c4b0
style Art fill:#ddd8e4,stroke:#b8b0c8
style Serve fill:#e8e0d4,stroke:#c8b89a
GPTQ and AWQ vs NF4/bitsandbytes
bitsandbytes NF4 quantization happens on the fly, right when a model is loaded, and is chosen because it's simple to set up during training. GPTQ and AWQ instead do a one-time, more careful calibration pass over a small dataset before the model is deployed - which costs extra setup time, but produces a quantized model that runs faster in production because it can use specialized fast INT4 GPU kernels that on-the-fly NF4 loading doesn't take advantage of the same way.
The distinction is about optimization target and when the quantization decision is made:
NF4 / bitsandbytes | GPTQ | AWQ | |
|---|---|---|---|
| When quantized | On-the-fly at from_pretrained load time | Offline, one-time calibration pass | Offline, one-time calibration pass |
| Optimized for | Fitting a frozen base model in memory during training (QLoRA) | Fast, low-memory inference | Fast, low-memory inference, better low-bit accuracy |
| Calibration data | None required | Few hundred representative samples (layer-wise reconstruction error minimization) | Few hundred representative samples (activation-aware channel scaling) |
| Inference kernel | Standard bitsandbytes dequant-on-the-fly kernels - not optimized for max serving throughput | Specialized INT4 GPU kernels (exllama, marlin) - fast | Specialized INT4 GPU kernels - fast, often better throughput/accuracy than GPTQ at equal bit-width |
| Typical use case | Training-time memory reduction (QLoRA) | Serving-time memory + latency reduction | Serving-time memory + latency reduction, preferred when accuracy at INT4 matters most |
How GPTQ works (briefly): quantizes weights layer by layer, using a calibration set to solve for the INT4 weight values that minimize the reconstruction error of that layer's output, correcting for quantization error in already-quantized weights as it proceeds through the layer.
How AWQ works (briefly): observes that a small fraction of weight channels are disproportionately important based on activation magnitudes (not weight magnitude), and preserves those channels at higher effective precision via per-channel scaling before quantizing everything to INT4 - this is why AWQ tends to retain accuracy better than GPTQ at the same bit-width on many benchmarks.
FP8 and FP4: The Default on Current GPUs
INT4 weight-only formats (GPTQ, AWQ) were the main serving option on A100-era GPUs. On Hopper (H100/H200) and Blackwell (B200/GB200) the tensor cores run FP8 - and on Blackwell FP4 - natively, which changes the default:
| Format | What is quantized | Why use it |
|---|---|---|
FP8 W8A8 (e.g. FP8_DYNAMIC, FP8_BLOCK) | Weights and activations to FP8 | ~Half the memory, faster prefill and decode on H100+, usually negligible quality loss - the default "safe" choice |
| INT4 weight-only (W4A16) via AWQ or GPTQ | Weights only | Smallest weights on any GPU; decode speed-up (memory-bound); activations stay 16-bit |
| NVFP4 / MXFP4 | Weights (and optionally activations) to 4-bit floats with block scales | Blackwell-native 4-bit; gpt-oss ships its MoE weights in MXFP4 |
| FP8 KV cache | The KV cache | Halves KV memory โ more concurrent sequences or longer context |
Offline quantization tooling has consolidated around llm-compressor (from the vLLM project), which implements FP8, GPTQ, AWQ, SmoothQuant and FP4 recipes and writes checkpoints vLLM loads directly. The older AutoAWQ package is no longer maintained; its functionality now lives in llm-compressor.
Code - Quantizing Your Own Fine-Tuned Model with llm-compressor
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "./my-merged-finetuned-model"
model = AutoModelForCausalLM.from_pretrained(model_path, dtype="auto")
tokenizer = AutoTokenizer.from_pretrained(model_path)
# FP8 weights + dynamic per-token FP8 activations: no calibration data needed
recipe = QuantizationModifier(targets="Linear", scheme="FP8_DYNAMIC", ignore=["lm_head"])
oneshot(model=model, recipe=recipe)
model.save_pretrained("./my-model-fp8")
tokenizer.save_pretrained("./my-model-fp8")
For INT4 weight-only, swap the recipe for an AWQ or GPTQ modifier (scheme W4A16) and pass a calibration dataset of a few hundred representative samples to oneshot.
Code - Serving the Quantized Checkpoint in vLLM
from vllm import LLM, SamplingParams
# vLLM reads the quantization format from the checkpoint's config - no flag needed
llm = LLM(model="./my-model-fp8", kv_cache_dtype="fp8")
print(llm.generate(["Explain KV caching in one sentence."], SamplingParams(max_tokens=64))[0].outputs[0].text)
# Or quantize a BF16 checkpoint to FP8 on the fly at load time (Hopper or newer):
# llm = LLM(model="Qwen/Qwen3-8B", quantization="fp8")
Accuracy and Speed Tradeoffs
Quantizing to INT4 typically costs a small amount of quality - usually small enough to be acceptable for most production tasks - in exchange for roughly a quarter of the memory footprint and meaningfully faster generation. Whether that trade is worth it depends entirely on how sensitive your task is to small quality regressions.
Rough, workload-dependent ranges (validate against your own eval set before shipping):
| Precision | Memory vs BF16 | Typical latency change | Typical quality delta |
|---|---|---|---|
| BF16/FP16 (baseline) | 1x | Baseline | Baseline |
| FP8 W8A8 (H100+) | ~50% | Faster prefill and decode | Usually negligible on large models |
| INT8 (W8A8 or weight-only) | ~50% | ~1.3-1.5x faster | Usually negligible |
| INT4 weight-only (AWQ/GPTQ) | ~25-30% | Faster decode (memory-bound); little prefill gain | Small, task-dependent - often 1-3% on standard benchmarks, more on precise numeric/reasoning output and on small models |
| FP4 (NVFP4/MXFP4, Blackwell) | ~25-30% | Fastest on Blackwell tensor cores | Small with good block scaling and calibration; validate on your tasks |
When to Quantize for Serving - and When Not To
Quantize for serving when GPU cost or throughput is the binding constraint and your task tolerates a small quality dip - customer support chat, summarization, classification-style generation. Skip it, or quantize less aggressively (INT8 instead of INT4), when the task is precision-sensitive - code generation, numeric reasoning, or anything where a small accuracy regression has an outsized downstream cost.
Quantize for serving when:
- GPU cost per request is the primary constraint (INT4 roughly quarters memory, often 2-3x's throughput)
- The task's eval metrics show acceptable degradation at the target bit-width on your own held-out set, not just published benchmarks
- You're deploying a single fixed model version at scale (calibration cost is a one-time investment, amortized over serving volume)
Don't quantize (or use INT8, not INT4) when:
- The task is precision-sensitive (code, math, structured extraction) where quantization-induced errors compound
- You're still iterating on the model (frequent re-quantization calibration cost outweighs the benefit)
- Serving volume is low enough that the memory/throughput win doesn't offset the quality risk and operational overhead of maintaining a quantized artifact alongside the full-precision one
A common production pattern: merge a LoRA/QLoRA fine-tune (see LoRA & QLoRA Hands-On) into a full-precision dense model first, then separately quantize that merged model for serving (FP8, AWQ or GPTQ via llm-compressor) - training-time and serving-time quantization decisions are made independently, at different stages of the pipeline.
Study Notes
Must-know for interviews:
- NF4/
bitsandbytes(training-time) and GPTQ/AWQ (serving-time) both quantize to low bit-width but optimize for different goals and use different kernels - GPTQ minimizes per-layer reconstruction error via calibration; AWQ preserves activation-important weight channels via per-channel scaling - AWQ often retains accuracy better at equal bit-width
- INT4 quantization typically costs ~25% of BF16 memory and 2-3x's throughput with optimized kernels, at a small task-dependent quality cost
- On H100-class GPUs, FP8 W8A8 (plus an FP8 KV cache) is the default safe serving quantization; FP4 (NVFP4/MXFP4) is the Blackwell-native 4-bit option
- llm-compressor is the current tool for producing FP8, AWQ, GPTQ and FP4 checkpoints for vLLM (AutoAWQ is no longer maintained)
- Quantize aggressively for cost-sensitive, quality-tolerant tasks; use FP8/INT8 or skip quantization for precision-sensitive tasks (code, math, structured extraction)
- A common production pipeline: merge LoRA adapter into a full-precision model, then separately AWQ/GPTQ-quantize the merged model for serving
Check Yourself
- On an H100 fleet, which quantization would you try first for serving?
- Why does weight-only 4-bit quantization speed up decode at small batch sizes?
- Why can't you just serve a
bitsandbytesNF4-quantized model the same way you'd serve a GPTQ/AWQ model? - What does AWQ preserve that plain uniform quantization doesn't?
- When would INT8 be the better choice over INT4 for a production deployment?
- A teammate proposes reusing the QLoRA NF4 config from training directly for production serving, to save setup time. What's the issue?
Exercises
Quantize a small instruct model to FP8 (or AWQ INT4) with llm-compressor, serve it with vLLM, and compare it with the BF16 original on 100 items of a task you care about and on decode throughput. Would you ship it?
Solution
Report accuracy with a paired comparison and throughput at 1 and 32 concurrent requests. FP8 is usually within noise of BF16 with a throughput gain on Hopper-class GPUs; INT4 gains more but may lose points on precise tasks - ship only if the quality delta is inside your tolerance.
References
- Frantar et al., GPTQ (2022); Lin et al., AWQ (2023); Xiao et al., SmoothQuant (2022)
- Micikevicius et al., FP8 Formats for Deep Learning (2022)
- vLLM project, llm-compressor and vLLM's quantization docs
- Dettmers et al., QLoRA (NF4) (2023)
Last reviewed: 2026-09