08 - Inference & Serving
Making models fast and cheap to serve: the KV cache and its memory math, vLLM's paged memory and continuous batching, quantized serving formats, the batch-size/latency trade-off and capacity planning, streaming and time-to-first-token, production budgets and monitoring, and the multi-node stack (disaggregated prefill/decode, KV-aware routing, speculative decoding, MoE and multi-LoRA serving).
- Compute KV-cache memory for a model and context, and explain how GQA, paging and prefix caching change it
- Configure and tune a vLLM deployment, and choose a serving engine and precision (BF16, FP8, INT4/FP4) for a GPU and task
- Plan capacity from an SLO using load tests - batch size vs concurrency, TTFT vs time per output token, goodput
- Serve a model behind a streaming endpoint and measure TTFT, latency percentiles and throughput under load
- Explain the modern serving stack - chunked prefill, disaggregation, KV-aware routing, speculative decoding, MoE and multi-LoRA serving
Chapter Map
| # | File | Topic | Difficulty |
|---|---|---|---|
| 1 | KV Cache & Inference | KV-cache memory math, MQA/GQA, PagedAttention, prefix caching, speculative decoding, continuous batching | Advanced |
| 2 | vLLM & Paged Attention | Why vLLM exists, PagedAttention KV-cache memory management, continuous vs static batching | Intermediate |
| 3 | Quantized Inference | FP8 and FP4 on current GPUs, AWQ/GPTQ INT4, llm-compressor, NF4 for training vs serving formats, when to quantize | Intermediate |
| 4 | Batching, Concurrency & Latency | Batch size vs concurrency, tokens/sec vs per-request latency, capacity planning basics | Intermediate |
| 5 | Streaming & TTFT | SSE/chunked streaming, TTFT vs total generation time, streaming's interaction with batching | Intermediate |
| 6 | Production Deployment | Latency budgets, caching hierarchy, context-window workarounds, cost, monitoring | Advanced |
| 7 | The Modern Serving Stack | Goodput and SLOs, chunked prefill, prefill/decode disaggregation, KV-aware routing and offloading, FP8/FP4, speculative decoding, MoE and multi-LoRA serving, Dynamo / llm-d | Advanced |
| 8 | Q&A Review Bank | 35 Q&A pairs across all topics in this module | All levels |
One home per concept: KV-cache math, GQA and speculative-decoding theory live in note 1; PagedAttention, prefix caching and the engine comparison in note 2; capacity planning in note 4; the multi-node stack in note 7. Note 6 is the production view and links to the others rather than repeating them.
Recommended Learning Paths
Path A: First Production Deployment, Start to Finish
- vLLM & Paged Attention - understand the serving engine before you configure it
- Batching, Concurrency & Latency - understand the throughput/latency trade-off you're tuning for
- Streaming & TTFT - decide whether your endpoint needs to stream
- FastAPI + vLLM Endpoint Code Lab - serve a real model end-to-end
- Quantized Inference - shrink the deployment footprint once it works
Path B: Interview Preparation (Accelerated)
- vLLM & Paged Attention - PagedAttention and continuous batching are asked constantly
- Batching, Concurrency & Latency - capacity planning math comes up in system design rounds
- Streaming & TTFT - TTFT vs throughput is a favorite follow-up
- Q&A Review Bank - drill all 35 questions
Path C: Cost & Capacity Engineering (Advanced)
- Quantized Inference - GPTQ/AWQ tradeoffs for shrinking GPU footprint
- Batching, Concurrency & Latency - how many concurrent users a given GPU can serve
- FastAPI + vLLM Endpoint Code Lab -
load_test.pyas a reusable capacity-planning template
Resources
- Q&A Review Bank - 35 Q&A pairs in this module
- Module quiz - every Check Yourself question in this module
- FastAPI + vLLM Endpoint Code Lab - hands-on streaming/non-streaming vLLM endpoint with a concurrent load test
Key Cross-References
- KV-cache fundamentals, GQA/MQA, and the theory behind Paged Attention → KV Cache & Inference - this module builds a hands-on serving stack on top of those foundations, it does not re-derive the memory math
- Conceptual overview of serving frameworks, latency budgets, and caching hierarchy → Production Deployment - this module goes deeper hands-on on the vLLM/FastAPI path specifically
- NF4/
bitsandbytestraining-time quantization and merge-vs-serve adapter decisions → Fine-Tuning Lab: LoRA & QLoRA Hands-On - contrast against this module's inference-time GPTQ/AWQ quantization - VRAM estimation and quantization formats (NF4, INT8, AWQ) → Pretraining: GPU & Hardware
Section Appendix
Summary & Key Terms - a quick recap of this section and its essential vocabulary.
Next Topic
Previous: 07 - Evaluation & Benchmarks · Next: 09 - Production Engineering