Contents
Map

08 ยท Inference & Serving

Quantized Inference

View as:

Quantized Inference

The One-Line Definition

Inference-time quantization (GPTQ, AWQ) and training-time quantization (NF4/bitsandbytes) both shrink model weights to fewer bits, but they optimize for different things - GPTQ/AWQ are calibrated once, offline, purely to make a frozen model cheaper and faster to serve, while NF4/bitsandbytes is designed to make a model cheap enough to train against, with speed as a secondary concern.

Learning objectives 40 min
By the end of this page you will be able to:
  • Distinguish training-time quantization (NF4) from serving formats (FP8, INT8, GPTQ, AWQ, FP4)
  • Choose a serving precision for a GPU generation and task sensitivity
  • Produce and serve a quantized checkpoint with llm-compressor and vLLM, and gate it on a quality evaluation

If the Fine-Tuning Lab (module 06) taught you to quantize a model so you could afford to fine-tune it on a single GPU, this page is about a related but distinct decision: quantizing an already-trained model purely to make it cheaper and faster to run for everyone who calls it in production. The tools and priorities are different even though the underlying idea - "use fewer bits per weight" - sounds the same.

This page assumes the quantization fundamentals (NF4 bin placement, VRAM formulas, INT8/INT4 tradeoffs) from Pretraining: GPU & Hardware and the QLoRA training-time setup from Fine-Tuning Lab: LoRA & QLoRA Hands-On. It is specifically about serving-side quantization - FP8 and FP4 on current GPUs, and INT4 weight-only formats (GPTQ, AWQ) - and when to reach for each.

flowchart LR
    FP["๐ŸŽฏ Full-precision\nfine-tuned model"]
    Cal["๐Ÿ“Š Calibration set\n(few hundred samples)"]
    GA["๐Ÿ”ข llm-compressor\nFP8 / AWQ / GPTQ / NVFP4\nquantize offline"]
    Art["๐Ÿ“ฆ Quantized checkpoint\n(FP8 โ‰ˆ 50%, INT4/FP4 โ‰ˆ 25-30%\nof BF16 size)"]
    Serve["๐Ÿš€ vLLM / SGLang serve\nwith quantized kernels"]

    FP --> Cal --> GA --> Art --> Serve

    style FP fill:#d8dfe8,stroke:#b0bac8
    style Cal fill:#e8e0d4,stroke:#c8b89a
    style GA fill:#dde4dc,stroke:#b0c4b0
    style Art fill:#ddd8e4,stroke:#b8b0c8
    style Serve fill:#e8e0d4,stroke:#c8b89a

GPTQ and AWQ vs NF4/bitsandbytes

bitsandbytes NF4 quantization happens on the fly, right when a model is loaded, and is chosen because it's simple to set up during training. GPTQ and AWQ instead do a one-time, more careful calibration pass over a small dataset before the model is deployed - which costs extra setup time, but produces a quantized model that runs faster in production because it can use specialized fast INT4 GPU kernels that on-the-fly NF4 loading doesn't take advantage of the same way.

The distinction is about optimization target and when the quantization decision is made:

NF4 / bitsandbytesGPTQAWQ
When quantizedOn-the-fly at from_pretrained load timeOffline, one-time calibration passOffline, one-time calibration pass
Optimized forFitting a frozen base model in memory during training (QLoRA)Fast, low-memory inferenceFast, low-memory inference, better low-bit accuracy
Calibration dataNone requiredFew hundred representative samples (layer-wise reconstruction error minimization)Few hundred representative samples (activation-aware channel scaling)
Inference kernelStandard bitsandbytes dequant-on-the-fly kernels - not optimized for max serving throughputSpecialized INT4 GPU kernels (exllama, marlin) - fastSpecialized INT4 GPU kernels - fast, often better throughput/accuracy than GPTQ at equal bit-width
Typical use caseTraining-time memory reduction (QLoRA)Serving-time memory + latency reductionServing-time memory + latency reduction, preferred when accuracy at INT4 matters most

How GPTQ works (briefly): quantizes weights layer by layer, using a calibration set to solve for the INT4 weight values that minimize the reconstruction error of that layer's output, correcting for quantization error in already-quantized weights as it proceeds through the layer.

How AWQ works (briefly): observes that a small fraction of weight channels are disproportionately important based on activation magnitudes (not weight magnitude), and preserves those channels at higher effective precision via per-channel scaling before quantizing everything to INT4 - this is why AWQ tends to retain accuracy better than GPTQ at the same bit-width on many benchmarks.

FP8 and FP4: The Default on Current GPUs

INT4 weight-only formats (GPTQ, AWQ) were the main serving option on A100-era GPUs. On Hopper (H100/H200) and Blackwell (B200/GB200) the tensor cores run FP8 - and on Blackwell FP4 - natively, which changes the default:

FormatWhat is quantizedWhy use it
FP8 W8A8 (e.g. FP8_DYNAMIC, FP8_BLOCK)Weights and activations to FP8~Half the memory, faster prefill and decode on H100+, usually negligible quality loss - the default "safe" choice
INT4 weight-only (W4A16) via AWQ or GPTQWeights onlySmallest weights on any GPU; decode speed-up (memory-bound); activations stay 16-bit
NVFP4 / MXFP4Weights (and optionally activations) to 4-bit floats with block scalesBlackwell-native 4-bit; gpt-oss ships its MoE weights in MXFP4
FP8 KV cacheThe KV cacheHalves KV memory โ†’ more concurrent sequences or longer context

Offline quantization tooling has consolidated around llm-compressor (from the vLLM project), which implements FP8, GPTQ, AWQ, SmoothQuant and FP4 recipes and writes checkpoints vLLM loads directly. The older AutoAWQ package is no longer maintained; its functionality now lives in llm-compressor.

Code - Quantizing Your Own Fine-Tuned Model with llm-compressor

from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
from transformers import AutoModelForCausalLM, AutoTokenizer

model_path = "./my-merged-finetuned-model"
model = AutoModelForCausalLM.from_pretrained(model_path, dtype="auto")
tokenizer = AutoTokenizer.from_pretrained(model_path)

# FP8 weights + dynamic per-token FP8 activations: no calibration data needed
recipe = QuantizationModifier(targets="Linear", scheme="FP8_DYNAMIC", ignore=["lm_head"])
oneshot(model=model, recipe=recipe)

model.save_pretrained("./my-model-fp8")
tokenizer.save_pretrained("./my-model-fp8")

For INT4 weight-only, swap the recipe for an AWQ or GPTQ modifier (scheme W4A16) and pass a calibration dataset of a few hundred representative samples to oneshot.

Code - Serving the Quantized Checkpoint in vLLM

from vllm import LLM, SamplingParams

# vLLM reads the quantization format from the checkpoint's config - no flag needed
llm = LLM(model="./my-model-fp8", kv_cache_dtype="fp8")
print(llm.generate(["Explain KV caching in one sentence."], SamplingParams(max_tokens=64))[0].outputs[0].text)

# Or quantize a BF16 checkpoint to FP8 on the fly at load time (Hopper or newer):
# llm = LLM(model="Qwen/Qwen3-8B", quantization="fp8")

Accuracy and Speed Tradeoffs

Quantizing to INT4 typically costs a small amount of quality - usually small enough to be acceptable for most production tasks - in exchange for roughly a quarter of the memory footprint and meaningfully faster generation. Whether that trade is worth it depends entirely on how sensitive your task is to small quality regressions.

Rough, workload-dependent ranges (validate against your own eval set before shipping):

PrecisionMemory vs BF16Typical latency changeTypical quality delta
BF16/FP16 (baseline)1xBaselineBaseline
FP8 W8A8 (H100+)~50%Faster prefill and decodeUsually negligible on large models
INT8 (W8A8 or weight-only)~50%~1.3-1.5x fasterUsually negligible
INT4 weight-only (AWQ/GPTQ)~25-30%Faster decode (memory-bound); little prefill gainSmall, task-dependent - often 1-3% on standard benchmarks, more on precise numeric/reasoning output and on small models
FP4 (NVFP4/MXFP4, Blackwell)~25-30%Fastest on Blackwell tensor coresSmall with good block scaling and calibration; validate on your tasks

When to Quantize for Serving - and When Not To

Quantize for serving when GPU cost or throughput is the binding constraint and your task tolerates a small quality dip - customer support chat, summarization, classification-style generation. Skip it, or quantize less aggressively (INT8 instead of INT4), when the task is precision-sensitive - code generation, numeric reasoning, or anything where a small accuracy regression has an outsized downstream cost.

Quantize for serving when:

  • GPU cost per request is the primary constraint (INT4 roughly quarters memory, often 2-3x's throughput)
  • The task's eval metrics show acceptable degradation at the target bit-width on your own held-out set, not just published benchmarks
  • You're deploying a single fixed model version at scale (calibration cost is a one-time investment, amortized over serving volume)

Don't quantize (or use INT8, not INT4) when:

  • The task is precision-sensitive (code, math, structured extraction) where quantization-induced errors compound
  • You're still iterating on the model (frequent re-quantization calibration cost outweighs the benefit)
  • Serving volume is low enough that the memory/throughput win doesn't offset the quality risk and operational overhead of maintaining a quantized artifact alongside the full-precision one

A common production pattern: merge a LoRA/QLoRA fine-tune (see LoRA & QLoRA Hands-On) into a full-precision dense model first, then separately quantize that merged model for serving (FP8, AWQ or GPTQ via llm-compressor) - training-time and serving-time quantization decisions are made independently, at different stages of the pipeline.


Study Notes

Must-know for interviews:

  • NF4/bitsandbytes (training-time) and GPTQ/AWQ (serving-time) both quantize to low bit-width but optimize for different goals and use different kernels
  • GPTQ minimizes per-layer reconstruction error via calibration; AWQ preserves activation-important weight channels via per-channel scaling - AWQ often retains accuracy better at equal bit-width
  • INT4 quantization typically costs ~25% of BF16 memory and 2-3x's throughput with optimized kernels, at a small task-dependent quality cost
  • On H100-class GPUs, FP8 W8A8 (plus an FP8 KV cache) is the default safe serving quantization; FP4 (NVFP4/MXFP4) is the Blackwell-native 4-bit option
  • llm-compressor is the current tool for producing FP8, AWQ, GPTQ and FP4 checkpoints for vLLM (AutoAWQ is no longer maintained)
  • Quantize aggressively for cost-sensitive, quality-tolerant tasks; use FP8/INT8 or skip quantization for precision-sensitive tasks (code, math, structured extraction)
  • A common production pipeline: merge LoRA adapter into a full-precision model, then separately AWQ/GPTQ-quantize the merged model for serving

Check Yourself

Check yourself
0 / 6 answered
  1. On an H100 fleet, which quantization would you try first for serving?
  2. Why does weight-only 4-bit quantization speed up decode at small batch sizes?
  3. Why can't you just serve a bitsandbytes NF4-quantized model the same way you'd serve a GPTQ/AWQ model?
  4. What does AWQ preserve that plain uniform quantization doesn't?
  5. When would INT8 be the better choice over INT4 for a production deployment?
  6. A teammate proposes reusing the QLoRA NF4 config from training directly for production serving, to save setup time. What's the issue?

Exercises

Exercise - Quantize and gate

Quantize a small instruct model to FP8 (or AWQ INT4) with llm-compressor, serve it with vLLM, and compare it with the BF16 original on 100 items of a task you care about and on decode throughput. Would you ship it?

Solution

Report accuracy with a paired comparison and throughput at 1 and 32 concurrent requests. FP8 is usually within noise of BF16 with a throughput gain on Hopper-class GPUs; INT4 gains more but may lose points on precise tasks - ship only if the quality delta is inside your tolerance.

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท