Contents
Map

08 ยท Inference & Serving

The Modern Serving Stack

View as:

The Modern Serving Stack

The earlier notes cover a single serving engine on one GPU or one server: PagedAttention, continuous batching, quantization, streaming. Serving at scale in 2025-2026 adds a layer above the engine - splitting prefill and decode onto different GPUs, routing requests to where their KV cache already lives, spreading mixture-of-experts models across many GPUs - plus newer precision formats and speculative decoding that actually ships. This note maps that stack.

Learning objectives 60 min
By the end of this page you will be able to:
  • Define TTFT, TPOT/ITL, throughput and goodput, and set SLOs in those terms
  • Explain chunked prefill and prefill/decode disaggregation, and when disaggregation pays off
  • Explain KV-cache-aware routing and KV offloading, and why they matter for multi-turn and agent traffic
  • Choose a precision (BF16, FP8, FP4) and KV-cache dtype for a deployment
  • Describe how speculative decoding (EAGLE-style draft heads, MTP, n-gram) and multi-LoRA serving are used in practice
  • Explain how large MoE models are served with expert parallelism
Prerequisites

Metrics That Define an SLO

MetricDefinitionWhat drives it
TTFT (time to first token)Request arrival โ†’ first output tokenQueueing + prefill of the prompt (compute-bound)
TPOT / ITL (time per output token / inter-token latency)Average gap between output tokensDecode step time (memory-bandwidth-bound) and batch size
E2E latencyTTFT + TPOT ร— (output tokens - 1)Both
ThroughputOutput tokens/s across all requestsBatch size, hardware, precision
GoodputRequests/s completed within the SLO (e.g. TTFT < 600 ms and TPOT < 40 ms at p99)The balance of all of the above

Throughput alone is misleading: a system can maximize tokens/s by batching aggressively while blowing through latency targets. Goodput (from the DistServe paper) is the number to optimize.

Worked example: with a 40 ms TPOT, a 500-token answer takes 20 s of decode after the first token - which is why streaming matters for chat, and why TPOT, not TTFT, dominates long-output workloads.


Prefill vs Decode: Chunking and Disaggregation

Concept

Prefill (process the whole prompt at once) is compute-bound; decode (one token per sequence per step) is memory-bound. Running both on the same GPU causes interference: a long prefill stalls every in-flight decode, spiking ITL.

Two fixes:

  • Chunked prefill (Sarathi-Serve): split long prefills into chunks and mix them with decode steps under a per-step token budget (vLLM's max_num_batched_tokens). Smooths ITL at a small TTFT cost. On by default in current engines.
  • Prefill/decode (P/D) disaggregation (DistServe, Splitwise, Mooncake): run prefill and decode on separate GPU pools, then transfer the KV cache from prefill to decode workers. Each pool is sized and parallelized for its own bottleneck, and neither interferes with the other.
flowchart LR
    U["๐Ÿ‘ค Requests"] --> R["๐Ÿงญ KV-aware router"]
    R --> PF["โšก Prefill pool<br/>compute-bound<br/>larger TP, big batches of prompt tokens"]
    PF -->|"KV cache transfer<br/>(NVLink / RDMA)"| DC["๐Ÿ” Decode pool<br/>memory-bound<br/>many concurrent sequences"]
    R -.->|"prefix already cached?<br/>skip / shorten prefill"| DC
    DC --> U2["๐Ÿ“ค Streamed tokens"]
    KVS[("๐Ÿ’พ KV cache tiers<br/>GPU โ†’ CPU RAM โ†’ SSD / remote")] <--> PF
    KVS <--> DC

    style PF fill:#e8e0d4,stroke:#c8b89a
    style DC fill:#d8dfe8,stroke:#b0bac8
    style R fill:#dde4dc,stroke:#b0c4b0
    style KVS fill:#e8e2d9,stroke:#ccc4b8

When disaggregation pays off: long prompts, strict ITL targets, and enough scale to run two pools. The cost is the KV transfer: an 8K-token prompt on a 70B GQA model is about 2.7 GB of KV cache - roughly 50 ms over a 400 Gb/s network link, or a few milliseconds over NVLink. For short prompts and small deployments, chunked prefill on unified workers is usually simpler and just as good.


KV-Cache-Aware Routing and Offloading

Agentic and multi-turn traffic resends a long, mostly unchanged prefix every turn (system prompt, tool definitions, conversation history). If each turn lands on a random replica, every turn recomputes that prefix.

  • KV-cache-aware routing sends a request to the replica that already holds the longest matching prefix, turning prefill into a cache hit. It is a core feature of NVIDIA Dynamo's router, llm-d's scheduler (built on the Kubernetes Gateway API Inference Extension) and SGLang's router.
  • KV offloading extends the cache beyond GPU memory - to CPU RAM, local SSD or a shared store - so evicted prefixes can be reloaded instead of recomputed (LMCache, Dynamo's KV block manager, Mooncake's KV-centric design).
  • Provider prompt caching is the API-side version of the same idea: cached input tokens are billed at a fraction of the normal price (see Prompt & Context Engineering).

Precision for Serving

ChoiceMemory vs BF16Typical quality impactHardware
BF16 weights1ร—BaselineAny modern GPU
FP8 weights + activations (W8A8)~0.5ร—Usually negligible on large modelsHopper (H100/H200) and newer
INT4 weight-only (AWQ, GPTQ)~0.25-0.3ร—Small; larger on small models and on math/codeBroad; good for memory-bound decode
FP4 (NVFP4, MXFP4)~0.25-0.3ร—Small with good calibration and block scalingBlackwell (native); gpt-oss ships MXFP4 experts
FP8 KV cacheKV cache ~0.5ร—Usually small; check long-context tasksHopper and newer

FP8 has become the default "safe" serving quantization on H100-class GPUs: vLLM can quantize weights to FP8 on the fly (--quantization fp8) or load pre-quantized checkpoints, and --kv-cache-dtype fp8 halves the KV cache. Offline quantization for serving has consolidated around llm-compressor (AWQ, GPTQ, FP8, NVFP4 recipes producing checkpoints vLLM loads directly). Details in Quantized Inference.


Speculative Decoding in Practice

Decode is memory-bound, so verifying several proposed tokens in one forward pass costs little more than generating one. Production options:

MethodDrafterNotes
EAGLE (1/2/3)A small head trained on the target model's hidden statesHigh acceptance rates; supported by vLLM, SGLang and TensorRT-LLM; pre-trained EAGLE heads exist for popular models
MTP headsMulti-token-prediction modules trained with the model (DeepSeek-V3)"Free" drafter - no separate training
n-gram / prompt lookupCopies spans from the promptNo model at all; great for editing, RAG and code where output repeats input
Draft modelA smaller model of the same familySimple; needs a compatible tokenizer

Gains are largest at low batch sizes (latency-bound serving) and shrink as the batch grows and the GPU becomes compute-bound. Measure TPOT at your real concurrency before and after enabling it.


Serving Mixture-of-Experts Models

A large MoE (DeepSeek-V3, Kimi K2, Qwen3-235B) has hundreds of experts but activates only a few per token. Serving it well means:

  • Expert parallelism (EP): spread experts across many GPUs; each token is dispatched (all-to-all) to the GPUs holding its experts and the results combined. Libraries such as DeepSeek's DeepEP optimize this communication.
  • Wide EP + data-parallel attention: attention runs data-parallel on each GPU while experts are sharded across a large EP group - the layout SGLang and vLLM use for DeepSeek-style models across dozens of GPUs, often combined with P/D disaggregation.
  • Big NVLink domains help: rack-scale domains such as GB200 NVL72 keep the all-to-all on NVLink instead of the network.
  • Load balancing: hot experts overload their GPUs; engines replicate popular experts (redundant experts) to spread load.

Multi-LoRA Serving

One base model, many fine-tuned adapters (per customer or per task): multi-LoRA serving keeps the base weights once in GPU memory and batches requests for different adapters together, applying each request's LoRA delta with specialized kernels (S-LoRA, Punica). In vLLM: start with --enable-lora and pass the adapter per request. It turns "one GPU per fine-tune" into "hundreds of fine-tunes per GPU".


The Orchestration Layer

ProjectWhat it adds on top of engines
NVIDIA DynamoDisaggregated serving, KV-aware router, KV block manager across memory tiers, NIXL for fast KV transfer; runs vLLM, SGLang or TensorRT-LLM workers
llm-dKubernetes-native distributed vLLM: KV-cache-aware scheduling via the Gateway API Inference Extension, P/D disaggregation, wide EP
SGLang router / PD modeCache-aware routing and prefill-decode disaggregation within the SGLang ecosystem
KServe, Ray ServeGeneral model-serving control planes (autoscaling, rollout) that host LLM engines - see Production Engineering

Check Yourself

Check yourself
0 / 4 answered
  1. A chat service must keep p99 inter-token latency under 50 ms, but long RAG prompts cause ITL spikes when their prefill runs. What is the first thing to try on a single-pool deployment?
  2. Why does KV-cache-aware routing matter most for agent workloads?
  3. Speculative decoding speeds up TPOT a lot at batch size 1 but barely at batch size 64. Why?
  4. What does goodput measure that throughput doesn't?

Exercises

Exercise - Size a disaggregated deployment

Traffic: 20 requests/s, prompts average 6K tokens, outputs average 300 tokens. SLO: TTFT p99 < 1.5 s, TPOT p99 < 40 ms. You measure that one prefill GPU (TP=1) processes about 30K prompt tokens/s at acceptable latency, and one decode GPU sustains about 60 concurrent sequences within the TPOT target.

  1. How many prefill GPUs do you need for the average load, before headroom?
  2. Roughly how many sequences are decoding concurrently at steady state (Little's law), and how many decode GPUs does that imply?
  3. What headroom would you add, and why?
Solution
  1. Prompt tokens/s = 20 ร— 6,000 = 120,000 โ†’ 120,000 / 30,000 = 4 prefill GPUs at average load.
  2. Each request decodes for about 300 ร— 40 ms = 12 s, so concurrency โ‰ˆ 20 req/s ร— 12 s = 240 sequences โ†’ 240 / 60 = 4 decode GPUs.
  3. Plan for peak rather than average (often 2-3ร— the mean) and for p99 rather than mean latency - e.g. 50-100% extra, with autoscaling on queue depth. Validate with a load test before committing.
Exercise - FP8 on your lab server

Run the FastAPI + vLLM lab on an H100-class GPU with and without --quantization fp8 --kv-cache-dtype fp8. Using load_test.py at 1, 16 and 64 concurrent users, report TTFT, tokens/s and max stable concurrency, then spot-check 20 outputs for quality differences.

Study Notes

Must-know:

  • SLO metrics: TTFT (prefill + queue), TPOT/ITL (decode), E2E, throughput, and goodput (requests/s within SLO)
  • Prefill is compute-bound, decode memory-bound; chunked prefill smooths ITL; P/D disaggregation separates them into pools at the cost of KV transfer
  • KV-aware routing + KV offloading turn repeated prefixes (agents, multi-turn) into cache hits; provider prompt caching is the API equivalent
  • FP8 W8A8 is the default safe serving precision on Hopper; FP4 (NVFP4/MXFP4) on Blackwell; FP8 KV cache halves cache memory; llm-compressor produces vLLM-ready checkpoints
  • Speculative decoding (EAGLE, MTP, n-gram) helps most at low batch sizes
  • Large MoE serving: expert parallelism with all-to-all, DP attention, redundant experts, big NVLink domains
  • Multi-LoRA serving batches many adapters over one base model
  • Orchestration: Dynamo, llm-d, SGLang router add disaggregation and cache-aware routing over engines

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท