vLLM and Paged Attention
The One-Line Definition
vLLM is the dominant open-source LLM serving engine because it treats the GPU's KV-cache memory the way an operating system treats virtual memory - paged, non-contiguous, and shared wherever possible - which is what lets it run continuous batching at near-full GPU utilization instead of the ~30-50% typical of naive serving.
- Explain the two problems vLLM solves - KV-cache fragmentation and idle GPU slots - and the mechanism for each
- Tune gpu_memory_utilization, block size, max_num_seqs, max_num_batched_tokens and max_model_len for a workload
- Explain how prefix caching reuses KV blocks across requests
- Choose a serving engine (vLLM, SGLang, TensorRT-LLM, llama.cpp) for a deployment
If you've only run a model in a notebook with model.generate(), this page explains the gap between that and a production endpoint serving hundreds of concurrent users. The short version: most of the engineering effort in serving an LLM well goes into not wasting GPU memory and not leaving the GPU idle, and vLLM's two headline features - Paged Attention and continuous batching - solve exactly those two problems.
This page assumes the KV-cache fundamentals (what the cache stores, why it exists, the memory formula) from KV Cache & Inference - it is not re-derived here. This page is specifically about how vLLM manages that cache in production and why the resulting throughput gain is so large.
Prerequisite: KV-cache memory math, MHA/MQA/GQA, and a first pass at Paged Attention and continuous batching are covered in KV Cache & Inference. This page goes deeper on the vLLM engine itself and how you configure/operate it.
flowchart LR
Req["๐ฅ Incoming requests"]
Sched["๐๏ธ vLLM scheduler\n(continuous batching)"]
Page["๐ PagedAttention\nKV-cache block manager"]
GPU["๐ฅ๏ธ GPU compute\n(near-100% utilization)"]
Out["๐ค Streamed tokens"]
Req --> Sched --> Page --> GPU --> Out
Sched -.evict finished /\ninsert waiting.-> Sched
style Req fill:#d8dfe8,stroke:#b0bac8
style Sched fill:#dde4dc,stroke:#b0c4b0
style Page fill:#e8e0d4,stroke:#c8b89a
style GPU fill:#ddd8e4,stroke:#b8b0c8
style Out fill:#d8dfe8,stroke:#b0bac8
Why vLLM Exists
Before vLLM (2023), most people serving open LLMs used the same code they trained or prototyped with - transformers' .generate() in a loop, one request at a time or in small fixed batches. That works for a demo. It falls over in production because GPU memory gets reserved wastefully and the GPU sits idle waiting for the slowest request in a batch to finish. vLLM was built specifically to fix both problems at once.
Two structural problems with naive serving:
- Memory fragmentation - a naive KV-cache allocator reserves a contiguous memory block per sequence sized for the maximum possible context length, even though most sequences are much shorter and grow gradually. Measured waste from this pattern: 60-80% of KV-cache memory unused.
- GPU idling under static batching - a fixed batch runs until its longest sequence finishes; every sequence that finished early leaves its GPU slot empty rather than being replaced.
vLLM's answer to each: Paged Attention (memory) and continuous batching (scheduling). Both are covered at the conceptual/math level in KV Cache & Inference; this page focuses on what that means operationally when you actually stand up a vLLM server.
Paged Attention in Practice
Instead of reserving one big continuous chunk of memory per conversation (most of which goes unused), vLLM splits GPU memory into small fixed-size blocks and hands them out on demand as a conversation grows - exactly like how an operating system pages a program's memory into RAM. When a conversation ends, its blocks go straight back into the pool for the next request, with almost no waste.
Concretely, vLLM's block_size (default 16 tokens) determines the granularity of memory allocation. Each sequence's logical KV cache is mapped to physical blocks via a per-sequence block table, exactly like a page table maps virtual to physical memory addresses. This has two practical consequences you configure around:
gpu_memory_utilization- fraction of GPU memory vLLM is allowed to claim for the KV-cache block pool (commonly0.85-0.95). Set too low and you under-utilize the GPU; too high and you risk OOM from other processes.- Automatic prefix caching (
enable_prefix_caching=True) - because blocks are addressed by content hash, identical prefixes (e.g. a shared system prompt) across different requests can literally point at the same physical blocks, with copy-on-write semantics if they diverge partway through. This is why vLLM's prefix caching is cheap to enable and high-ROI for chatbot/RAG workloads with repeated system prompts.
flowchart TD
subgraph Logical["Logical KV cache (per request)"]
L1["Seq A: tokens 0-63"]
L2["Seq B: tokens 0-31"]
end
subgraph Physical["Physical GPU blocks (16 tokens each)"]
B1["Block 3"]
B2["Block 7"]
B3["Block 12"]
B4["Block 19"]
B5["Block 1"]
B6["Block 4"]
end
L1 -->|block table A| B1
L1 --> B2
L1 --> B3
L1 --> B4
L2 -->|block table B| B5
L2 --> B6
style Logical fill:#d8dfe8,stroke:#b0bac8
style Physical fill:#e8e0d4,stroke:#c8b89a
Measured Results
| Metric | Naive contiguous allocation | Paged Attention |
|---|---|---|
| Memory waste from fragmentation | 60-80% | < 4% |
| Sequences served concurrently (same GPU) | Baseline | 2-4x more |
| Throughput | Baseline | 2-4x higher |
Continuous Batching in Practice
vLLM never waits for a whole batch to finish together. After every single decode step, it checks which sequences in the current batch are done, evicts them immediately, and slots in the next waiting request - so the GPU is always working on a full batch, not draining down to a handful of stragglers.
This is scheduled by vLLM's continuous (in-flight) batching engine, distinct from the "dynamic batching" some frameworks do at the request-queue level (grouping same-length requests before a fixed batch starts). vLLM operates at the token-generation-step granularity, which is what produces the 5-10x throughput gain over static batching referenced in KV Cache & Inference.
Key scheduler knobs:
| Parameter | Controls |
|---|---|
max_num_seqs | Maximum number of sequences batched together per step |
max_num_batched_tokens | Token budget per scheduling step - caps prefill from monopolizing a step and starving decode-phase sequences (chunked prefill) |
max_model_len | Maximum context length vLLM will allocate KV-cache blocks for |
Serving Framework Comparison
| Engine | Strengths | Typical use |
|---|---|---|
| vLLM | PagedAttention, automatic prefix caching, broad model and hardware support, OpenAI-compatible server (vllm serve), speculative decoding, multi-LoRA, FP8/FP4 quantized serving. V1 engine only since v0.11 | The default open-source serving engine |
| SGLang | RadixAttention (tree-structured prefix cache), fast structured output, strong MoE / expert-parallel and prefill-decode disaggregation support | Heavy prefix reuse (agents, few-shot), large MoE models such as DeepSeek-V3 |
| TensorRT-LLM | NVIDIA-optimized kernels and in-flight batching, FP8/NVFP4 on Hopper and Blackwell | Maximum throughput on NVIDIA GPUs, often behind NVIDIA Dynamo or Triton |
| TGI (Hugging Face) | Continuous batching, simple deployment of Hub models | Hugging Face-centric deployments |
| llama.cpp / Ollama | GGUF quantized models on CPUs, Apple silicon and consumer GPUs | Local and edge inference, prototyping |
Above the engine sits an orchestration layer for multi-node deployments - NVIDIA Dynamo, llm-d, and SGLang's own router - which adds disaggregated prefill/decode and KV-cache-aware request routing. See The Modern Serving Stack.
Study Notes
Must-know for interviews:
- vLLM's two headline features are Paged Attention (KV-cache memory) and continuous batching (scheduling) - they solve memory fragmentation and GPU idling respectively
- Paged Attention allocates fixed-size blocks (default 16 tokens) on demand via a per-sequence block table, borrowing the OS virtual-memory-paging model
- Automatic prefix caching works because content-addressed blocks let identical prefixes across requests share physical memory with copy-on-write
- Continuous batching evicts finished sequences and inserts waiting ones after every decode step, not at fixed batch boundaries
gpu_memory_utilization,max_num_seqs, andmax_num_batched_tokensare the primary vLLM server tuning knobs- Combined, these features typically deliver 2-4x higher throughput than naive serving on identical hardware
Check Yourself
- What did the vLLM paper measure as KV-cache memory waste in earlier serving systems versus PagedAttention?
- A long prompt arrives while 50 sequences are decoding. Which setting stops its prefill from stalling their decode steps?
- What does vLLM's
block_sizecontrol? - Why is chunked prefill needed alongside continuous batching?
- How does prefix caching avoid recomputing a shared system prompt for every request?
- What's the practical effect of setting
gpu_memory_utilizationtoo low? - Why does naive
transformers.generate()in a loop not scale to production traffic even on a large GPU?
Exercises
An 8B model with 32 layers, 8 KV heads of dimension 128, served in BF16 on an 80 GB GPU with gpu_memory_utilization = 0.9. How many 16-token KV blocks fit after the ~16 GB of weights, and how many concurrent 4K-token sequences is that?
Solution
KV bytes per token = 2 ร 32 ร 8 ร 128 ร 2 = 131,072 (128 KiB). Budget โ 0.9 ร 80 - 16 = 56 GB (ignoring activations and CUDA graphs), so about 56e9 / 131,072 โ 427,000 tokens โ 26,700 blocks of 16 tokens - roughly 104 sequences at 4,096 tokens each. Real numbers are lower because of activation memory and fragmentation at the last block of each sequence.
References
- Kwon et al., Efficient Memory Management for LLM Serving with PagedAttention (2023)
- vLLM project, vLLM documentation (2026)
- Zheng et al., SGLang (2024)
Last reviewed: 2026-09