Contents
Map

08 · Inference & Serving

Overview

View as:

08 - Inference & Serving

Making models fast and cheap to serve: the KV cache and its memory math, vLLM's paged memory and continuous batching, quantized serving formats, the batch-size/latency trade-off and capacity planning, streaming and time-to-first-token, production budgets and monitoring, and the multi-node stack (disaggregated prefill/decode, KV-aware routing, speculative decoding, MoE and multi-LoRA serving).

Learning objectives 9-11 hours (notes + lab)
By the end of this module you will be able to:
  • Compute KV-cache memory for a model and context, and explain how GQA, paging and prefix caching change it
  • Configure and tune a vLLM deployment, and choose a serving engine and precision (BF16, FP8, INT4/FP4) for a GPU and task
  • Plan capacity from an SLO using load tests - batch size vs concurrency, TTFT vs time per output token, goodput
  • Serve a model behind a streaming endpoint and measure TTFT, latency percentiles and throughput under load
  • Explain the modern serving stack - chunked prefill, disaggregation, KV-aware routing, speculative decoding, MoE and multi-LoRA serving

Chapter Map

#FileTopicDifficulty
1KV Cache & InferenceKV-cache memory math, MQA/GQA, PagedAttention, prefix caching, speculative decoding, continuous batchingAdvanced
2vLLM & Paged AttentionWhy vLLM exists, PagedAttention KV-cache memory management, continuous vs static batchingIntermediate
3Quantized InferenceFP8 and FP4 on current GPUs, AWQ/GPTQ INT4, llm-compressor, NF4 for training vs serving formats, when to quantizeIntermediate
4Batching, Concurrency & LatencyBatch size vs concurrency, tokens/sec vs per-request latency, capacity planning basicsIntermediate
5Streaming & TTFTSSE/chunked streaming, TTFT vs total generation time, streaming's interaction with batchingIntermediate
6Production DeploymentLatency budgets, caching hierarchy, context-window workarounds, cost, monitoringAdvanced
7The Modern Serving StackGoodput and SLOs, chunked prefill, prefill/decode disaggregation, KV-aware routing and offloading, FP8/FP4, speculative decoding, MoE and multi-LoRA serving, Dynamo / llm-dAdvanced
8Q&A Review Bank35 Q&A pairs across all topics in this moduleAll levels

One home per concept: KV-cache math, GQA and speculative-decoding theory live in note 1; PagedAttention, prefix caching and the engine comparison in note 2; capacity planning in note 4; the multi-node stack in note 7. Note 6 is the production view and links to the others rather than repeating them.

Path A: First Production Deployment, Start to Finish

  1. vLLM & Paged Attention - understand the serving engine before you configure it
  2. Batching, Concurrency & Latency - understand the throughput/latency trade-off you're tuning for
  3. Streaming & TTFT - decide whether your endpoint needs to stream
  4. FastAPI + vLLM Endpoint Code Lab - serve a real model end-to-end
  5. Quantized Inference - shrink the deployment footprint once it works

Path B: Interview Preparation (Accelerated)

  1. vLLM & Paged Attention - PagedAttention and continuous batching are asked constantly
  2. Batching, Concurrency & Latency - capacity planning math comes up in system design rounds
  3. Streaming & TTFT - TTFT vs throughput is a favorite follow-up
  4. Q&A Review Bank - drill all 35 questions

Path C: Cost & Capacity Engineering (Advanced)

  1. Quantized Inference - GPTQ/AWQ tradeoffs for shrinking GPU footprint
  2. Batching, Concurrency & Latency - how many concurrent users a given GPU can serve
  3. FastAPI + vLLM Endpoint Code Lab - load_test.py as a reusable capacity-planning template

Resources

Key Cross-References

  • KV-cache fundamentals, GQA/MQA, and the theory behind Paged Attention → KV Cache & Inference - this module builds a hands-on serving stack on top of those foundations, it does not re-derive the memory math
  • Conceptual overview of serving frameworks, latency budgets, and caching hierarchy → Production Deployment - this module goes deeper hands-on on the vLLM/FastAPI path specifically
  • NF4/bitsandbytes training-time quantization and merge-vs-serve adapter decisions → Fine-Tuning Lab: LoRA & QLoRA Hands-On - contrast against this module's inference-time GPTQ/AWQ quantization
  • VRAM estimation and quantization formats (NF4, INT8, AWQ) → Pretraining: GPU & Hardware

Section Appendix

Summary & Key Terms - a quick recap of this section and its essential vocabulary.


Next Topic

Previous: 07 - Evaluation & Benchmarks · Next: 09 - Production Engineering

⚡AI-assisted content - always verify, always explore multiple perspectives·