Contents
Map

09 ยท Production Engineering

LLM Serving on Kubernetes

View as:

LLM Serving on Kubernetes

Kubernetes & Helm covers the core objects and GPU node pools. Running LLMs on Kubernetes at scale adds problems generic web services don't have: GPUs that must be shared or scheduled by attributes, models too big for one node, autoscaling that CPU metrics can't drive, and cold starts measured in minutes. This note maps the Kubernetes-native tools that address them.

Learning objectives 50 min
By the end of this page you will be able to:
  • Choose between whole-GPU allocation, MIG, time-slicing and DRA-based scheduling for a workload
  • Describe the serving layer options - KServe, LeaderWorkerSet for multi-node models, llm-d and the Gateway API Inference Extension
  • Design autoscaling on the right signal (queue depth, KV-cache usage) and explain why GPU utilization is a poor one
  • Estimate and reduce cold-start time for large models

The Stack at a Glance

flowchart TD
    U["๐Ÿ‘ค Clients"] --> GW["๐Ÿšช Gateway API +<br/>Inference Extension<br/>(model- and KV-cache-aware routing)"]
    GW --> POOL["๐Ÿงฉ Inference pool<br/>vLLM / SGLang pods"]
    POOL --> AS["๐Ÿ“ˆ Autoscaler<br/>KEDA / HPA on queue depth"]
    POOL --> LWS["๐Ÿ”— LeaderWorkerSet<br/>one model across several nodes"]
    POOL --> GPU["๐Ÿ–ฅ๏ธ GPU nodes<br/>GPU Operator ยท MIG / time-slicing ยท DRA"]
    CTRL["๐ŸŽ›๏ธ KServe / llm-d<br/>control plane"] -.-> POOL
    CTRL -.-> GW

    style GW fill:#d8dfe8,stroke:#b0bac8
    style POOL fill:#dde4dc,stroke:#b0c4b0
    style GPU fill:#e8e0d4,stroke:#c8b89a
    style CTRL fill:#ddd8e4,stroke:#b8b0c8

GPUs as a Kubernetes Resource

Getting GPUs onto nodes

The NVIDIA GPU Operator installs and manages everything a node needs - driver, container toolkit, device plugin, DCGM exporter for metrics, MIG manager - as Kubernetes-managed components, so GPU nodes are configured declaratively rather than by hand.

Sharing and allocating GPUs

ApproachHow it worksIsolationUse when
Whole GPU (nvidia.com/gpu: 1)Device plugin hands out integer GPUsFullMost LLM serving - models need the whole GPU's memory
MIG (A100/H100 and newer)Partition one GPU into up to 7 hardware-isolated instances with their own memory and computeHardwareMany small models (embeddings, rerankers, small LLMs) on one GPU
Time-slicingSeveral pods share a GPU in turnsNone - memory is shared, no fault isolationDev/test, bursty low-priority workloads
DRA (Dynamic Resource Allocation)Pods request devices by attributes (memory size, model, interconnect) through ResourceClaims; GA since Kubernetes 1.34Depends on driverHeterogeneous fleets; requesting "an H100 with โ‰ฅ 80 GB" or topology-aware placement instead of an opaque count

The Serving Layer

ProjectWhat it does
KServeA model-serving control plane (CNCF): InferenceService resources with autoscaling, canary rollout and standard protocols; supports vLLM as the LLM runtime
LeaderWorkerSet (LWS)Kubernetes API for a group of pods that together serve one replica - a leader plus workers across nodes - for models that need multi-node tensor/pipeline parallelism; scaled and restarted as a unit
Gateway API Inference ExtensionExtends Kubernetes Gateway API with InferencePool and model-aware routing, so the gateway can pick a replica by queue length, KV-cache utilization or loaded LoRA adapter instead of round-robin
llm-dA Kubernetes-native distributed-inference stack on vLLM: builds on the Inference Extension for KV-cache-aware scheduling, and adds prefill/decode disaggregation and wide expert parallelism
NVIDIA DynamoDistributed-inference framework with its own Kubernetes operator (see The Modern Serving Stack)

Why round-robin load balancing fails for LLMs: requests differ by 100ร— in cost (a 50-token question vs a 100K-token document), and replicas differ in what they have cached. A replica with a hot prefix cache or a short queue is worth far more than "the next one in the list".


Autoscaling on the Right Signal

GPU utilization is a poor scaling signal: a vLLM server batching aggressively reports near-100% utilization whether it has 5 or 500 requests waiting.

SignalWhy it's useful
Waiting requests (vllm:num_requests_waiting)Direct measure of queueing โ†’ TTFT risk
KV-cache usage (vllm:kv_cache_usage_perc)Near 100% means preemptions and recomputation are coming
TTFT / ITL p95 (vllm:time_to_first_token_seconds, vllm:inter_token_latency_seconds)Scale on the SLO itself (reacts later, but is what users feel)

KEDA scales Deployments on Prometheus queries (and scales to zero for idle models); HPA can do the same through a custom-metrics adapter. Set scale-up thresholds early enough to absorb the cold-start time below.


Cold Starts

A new replica must pull a multi-GB image, load weights into GPU memory, and warm up (CUDA graphs, compilation). For large models this dominates scale-up time.

Weights to load รท read bandwidth โ‰ˆ load time
  70 GB (a 70B model in FP8) from object storage at ~1 GB/s  โ†’  ~70 s
  same from local NVMe at ~5 GB/s                             โ†’  ~14 s
  plus image pull, engine start-up, CUDA-graph capture and warm-up

Ways to cut it:

  • Keep weights out of the image and on fast storage close to the GPU - local NVMe caches, pre-populated persistent volumes, or node-level caches.
  • Stream weights directly into GPU memory rather than download-then-load (e.g. Run:ai Model Streamer, supported by vLLM).
  • Pre-pull images on GPU nodes; keep images slim (multi-stage builds - see Docker for GPU Inference).
  • Keep warm capacity (minimum replicas, or scale-to-zero only for rarely used models).
  • Use smaller or quantized checkpoints where quality allows - FP8 halves the bytes to load.

Check Yourself

Check yourself
0 / 3 answered
  1. A vLLM deployment shows 98% GPU utilization at both 3 AM and peak hour, but TTFT at peak is 10ร— worse. What should autoscaling watch instead?
  2. You need to serve 12 small embedding and reranking models with strict isolation on a few H100s. Which GPU-sharing approach fits best?
  3. What problem does LeaderWorkerSet solve that a Deployment doesn't?

Exercises

Exercise - Design a scaling policy

Your 70B FP8 model runs on 2ร— H100 per replica. Cold start is about 3 minutes; traffic doubles over roughly 10 minutes at the morning peak; the SLO is TTFT p95 under 2 s. Propose the metric, thresholds, minimum replicas and scale-down behaviour for KEDA, and justify each against the cold-start time.

Exercise - Map the course lab to Kubernetes

Take the Helm chart lab. List what you would change to (a) scale on vllm:num_requests_waiting via KEDA instead of GPU utilization, (b) load weights from a pre-populated volume, and (c) put the service behind a Gateway API route.

Study Notes

Must-know:

  • GPU Operator manages drivers, toolkit, device plugin, DCGM and MIG on nodes
  • GPU sharing: whole GPU (default for LLMs), MIG (hardware isolation), time-slicing (no isolation), DRA (attribute-based claims, GA in Kubernetes 1.34)
  • Serving layer: KServe (control plane), LeaderWorkerSet (multi-node replicas), Gateway API Inference Extension (model- and cache-aware routing), llm-d (distributed vLLM)
  • Autoscale on queue depth, KV-cache usage or latency SLOs - not GPU utilization; KEDA supports Prometheus signals and scale-to-zero
  • Cold start โ‰ˆ weights รท read bandwidth + start-up; keep weights on fast local storage, stream them, pre-pull images, keep warm capacity

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท