Contents
Map

08 ยท Inference & Serving

vLLM & Paged Attention

View as:

vLLM and Paged Attention

The One-Line Definition

vLLM is the dominant open-source LLM serving engine because it treats the GPU's KV-cache memory the way an operating system treats virtual memory - paged, non-contiguous, and shared wherever possible - which is what lets it run continuous batching at near-full GPU utilization instead of the ~30-50% typical of naive serving.

Learning objectives 40 min
By the end of this page you will be able to:
  • Explain the two problems vLLM solves - KV-cache fragmentation and idle GPU slots - and the mechanism for each
  • Tune gpu_memory_utilization, block size, max_num_seqs, max_num_batched_tokens and max_model_len for a workload
  • Explain how prefix caching reuses KV blocks across requests
  • Choose a serving engine (vLLM, SGLang, TensorRT-LLM, llama.cpp) for a deployment

If you've only run a model in a notebook with model.generate(), this page explains the gap between that and a production endpoint serving hundreds of concurrent users. The short version: most of the engineering effort in serving an LLM well goes into not wasting GPU memory and not leaving the GPU idle, and vLLM's two headline features - Paged Attention and continuous batching - solve exactly those two problems.

This page assumes the KV-cache fundamentals (what the cache stores, why it exists, the memory formula) from KV Cache & Inference - it is not re-derived here. This page is specifically about how vLLM manages that cache in production and why the resulting throughput gain is so large.

Prerequisite: KV-cache memory math, MHA/MQA/GQA, and a first pass at Paged Attention and continuous batching are covered in KV Cache & Inference. This page goes deeper on the vLLM engine itself and how you configure/operate it.

flowchart LR
    Req["๐Ÿ“ฅ Incoming requests"]
    Sched["๐Ÿ—‚๏ธ vLLM scheduler\n(continuous batching)"]
    Page["๐Ÿ“„ PagedAttention\nKV-cache block manager"]
    GPU["๐Ÿ–ฅ๏ธ GPU compute\n(near-100% utilization)"]
    Out["๐Ÿ“ค Streamed tokens"]

    Req --> Sched --> Page --> GPU --> Out
    Sched -.evict finished /\ninsert waiting.-> Sched

    style Req fill:#d8dfe8,stroke:#b0bac8
    style Sched fill:#dde4dc,stroke:#b0c4b0
    style Page fill:#e8e0d4,stroke:#c8b89a
    style GPU fill:#ddd8e4,stroke:#b8b0c8
    style Out fill:#d8dfe8,stroke:#b0bac8

Why vLLM Exists

Before vLLM (2023), most people serving open LLMs used the same code they trained or prototyped with - transformers' .generate() in a loop, one request at a time or in small fixed batches. That works for a demo. It falls over in production because GPU memory gets reserved wastefully and the GPU sits idle waiting for the slowest request in a batch to finish. vLLM was built specifically to fix both problems at once.

Two structural problems with naive serving:

  1. Memory fragmentation - a naive KV-cache allocator reserves a contiguous memory block per sequence sized for the maximum possible context length, even though most sequences are much shorter and grow gradually. Measured waste from this pattern: 60-80% of KV-cache memory unused.
  2. GPU idling under static batching - a fixed batch runs until its longest sequence finishes; every sequence that finished early leaves its GPU slot empty rather than being replaced.

vLLM's answer to each: Paged Attention (memory) and continuous batching (scheduling). Both are covered at the conceptual/math level in KV Cache & Inference; this page focuses on what that means operationally when you actually stand up a vLLM server.


Paged Attention in Practice

Instead of reserving one big continuous chunk of memory per conversation (most of which goes unused), vLLM splits GPU memory into small fixed-size blocks and hands them out on demand as a conversation grows - exactly like how an operating system pages a program's memory into RAM. When a conversation ends, its blocks go straight back into the pool for the next request, with almost no waste.

Concretely, vLLM's block_size (default 16 tokens) determines the granularity of memory allocation. Each sequence's logical KV cache is mapped to physical blocks via a per-sequence block table, exactly like a page table maps virtual to physical memory addresses. This has two practical consequences you configure around:

  • gpu_memory_utilization - fraction of GPU memory vLLM is allowed to claim for the KV-cache block pool (commonly 0.85-0.95). Set too low and you under-utilize the GPU; too high and you risk OOM from other processes.
  • Automatic prefix caching (enable_prefix_caching=True) - because blocks are addressed by content hash, identical prefixes (e.g. a shared system prompt) across different requests can literally point at the same physical blocks, with copy-on-write semantics if they diverge partway through. This is why vLLM's prefix caching is cheap to enable and high-ROI for chatbot/RAG workloads with repeated system prompts.
flowchart TD
    subgraph Logical["Logical KV cache (per request)"]
        L1["Seq A: tokens 0-63"]
        L2["Seq B: tokens 0-31"]
    end
    subgraph Physical["Physical GPU blocks (16 tokens each)"]
        B1["Block 3"]
        B2["Block 7"]
        B3["Block 12"]
        B4["Block 19"]
        B5["Block 1"]
        B6["Block 4"]
    end
    L1 -->|block table A| B1
    L1 --> B2
    L1 --> B3
    L1 --> B4
    L2 -->|block table B| B5
    L2 --> B6

    style Logical fill:#d8dfe8,stroke:#b0bac8
    style Physical fill:#e8e0d4,stroke:#c8b89a

Measured Results

MetricNaive contiguous allocationPaged Attention
Memory waste from fragmentation60-80%< 4%
Sequences served concurrently (same GPU)Baseline2-4x more
ThroughputBaseline2-4x higher

Continuous Batching in Practice

vLLM never waits for a whole batch to finish together. After every single decode step, it checks which sequences in the current batch are done, evicts them immediately, and slots in the next waiting request - so the GPU is always working on a full batch, not draining down to a handful of stragglers.

This is scheduled by vLLM's continuous (in-flight) batching engine, distinct from the "dynamic batching" some frameworks do at the request-queue level (grouping same-length requests before a fixed batch starts). vLLM operates at the token-generation-step granularity, which is what produces the 5-10x throughput gain over static batching referenced in KV Cache & Inference.

Key scheduler knobs:

ParameterControls
max_num_seqsMaximum number of sequences batched together per step
max_num_batched_tokensToken budget per scheduling step - caps prefill from monopolizing a step and starving decode-phase sequences (chunked prefill)
max_model_lenMaximum context length vLLM will allocate KV-cache blocks for

Serving Framework Comparison

EngineStrengthsTypical use
vLLMPagedAttention, automatic prefix caching, broad model and hardware support, OpenAI-compatible server (vllm serve), speculative decoding, multi-LoRA, FP8/FP4 quantized serving. V1 engine only since v0.11The default open-source serving engine
SGLangRadixAttention (tree-structured prefix cache), fast structured output, strong MoE / expert-parallel and prefill-decode disaggregation supportHeavy prefix reuse (agents, few-shot), large MoE models such as DeepSeek-V3
TensorRT-LLMNVIDIA-optimized kernels and in-flight batching, FP8/NVFP4 on Hopper and BlackwellMaximum throughput on NVIDIA GPUs, often behind NVIDIA Dynamo or Triton
TGI (Hugging Face)Continuous batching, simple deployment of Hub modelsHugging Face-centric deployments
llama.cpp / OllamaGGUF quantized models on CPUs, Apple silicon and consumer GPUsLocal and edge inference, prototyping

Above the engine sits an orchestration layer for multi-node deployments - NVIDIA Dynamo, llm-d, and SGLang's own router - which adds disaggregated prefill/decode and KV-cache-aware request routing. See The Modern Serving Stack.


Study Notes

Must-know for interviews:

  • vLLM's two headline features are Paged Attention (KV-cache memory) and continuous batching (scheduling) - they solve memory fragmentation and GPU idling respectively
  • Paged Attention allocates fixed-size blocks (default 16 tokens) on demand via a per-sequence block table, borrowing the OS virtual-memory-paging model
  • Automatic prefix caching works because content-addressed blocks let identical prefixes across requests share physical memory with copy-on-write
  • Continuous batching evicts finished sequences and inserts waiting ones after every decode step, not at fixed batch boundaries
  • gpu_memory_utilization, max_num_seqs, and max_num_batched_tokens are the primary vLLM server tuning knobs
  • Combined, these features typically deliver 2-4x higher throughput than naive serving on identical hardware

Check Yourself

Check yourself
0 / 7 answered
  1. What did the vLLM paper measure as KV-cache memory waste in earlier serving systems versus PagedAttention?
  2. A long prompt arrives while 50 sequences are decoding. Which setting stops its prefill from stalling their decode steps?
  3. What does vLLM's block_size control?
  4. Why is chunked prefill needed alongside continuous batching?
  5. How does prefix caching avoid recomputing a shared system prompt for every request?
  6. What's the practical effect of setting gpu_memory_utilization too low?
  7. Why does naive transformers.generate() in a loop not scale to production traffic even on a large GPU?

Exercises

Exercise - Size the block pool

An 8B model with 32 layers, 8 KV heads of dimension 128, served in BF16 on an 80 GB GPU with gpu_memory_utilization = 0.9. How many 16-token KV blocks fit after the ~16 GB of weights, and how many concurrent 4K-token sequences is that?

Solution

KV bytes per token = 2 ร— 32 ร— 8 ร— 128 ร— 2 = 131,072 (128 KiB). Budget โ‰ˆ 0.9 ร— 80 - 16 = 56 GB (ignoring activations and CUDA graphs), so about 56e9 / 131,072 โ‰ˆ 427,000 tokens โ‰ˆ 26,700 blocks of 16 tokens - roughly 104 sequences at 4,096 tokens each. Real numbers are lower because of activation memory and fragmentation at the last block of each sequence.

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท