Contents
Map

09 ยท Production Engineering

Kubernetes & Helm

View as:

Kubernetes and Helm

Kubernetes is the system that keeps your model-serving app running in production - it restarts crashed pods, spreads traffic across replicas, and scales up when demand spikes. Helm is the packaging layer on top: instead of hand-editing a dozen YAML files per environment, you fill in one small values file and Helm generates the rest.

A GPU inference workload on Kubernetes needs the same core objects as any web service (Deployment, Service, Ingress) plus GPU-specific scheduling concerns (node pools, resource requests, autoscaling on serving signals such as queue depth rather than CPU). Helm templates the boilerplate so the same chart deploys to dev/staging/prod with different values.yaml overrides.

Learning objectives 35 min
By the end of this page you will be able to:
  • Name the four Kubernetes objects behind most model-serving deployments and what each does
  • Schedule GPU pods correctly with nvidia.com/gpu limits, node-pool taints and tolerations
  • Explain why the default HPA cannot scale an LLM server well and what signal to use instead
  • Structure a Helm chart and choose between helm template, --dry-run and upgrade --install

The Core Objects

Four building blocks cover most of what you need: a Deployment (how many copies of your app to run and how to update them), a Service (a stable network address for those copies), an Ingress (the public door into your cluster), and an HPA (a rule that adds more copies automatically under load).

  • Deployment - declares desired replica count, the pod template (container image, resource requests/limits, GPU count via nvidia.com/gpu: 1), and the rollout strategy (RollingUpdate with maxSurge/maxUnavailable)
  • Service - a stable ClusterIP/virtual address in front of a set of pods selected by label, so pods can restart or reschedule without clients needing to track individual pod IPs
  • Ingress - routes external HTTP(S) traffic into the cluster to the right Service, usually via an ingress controller (nginx, Traefik) that terminates TLS
  • HorizontalPodAutoscaler (HPA) - watches a metric (CPU/memory by default; anything else, such as GPU utilization from the DCGM exporter or vLLM's queue depth, via a custom-metrics adapter) and adjusts replica count within a min/max band. For LLM serving, queue depth is the better signal - see LLM Serving on Kubernetes
flowchart LR
    U["๐ŸŒ Client Request"] --> I["๐Ÿšช Ingress\nTLS termination,\nrouting"]
    I --> S["๐Ÿ”Œ Service\nStable address,\nload balances pods"]
    S --> P1["๐Ÿ“ฆ Pod 1\nGPU: 1"]
    S --> P2["๐Ÿ“ฆ Pod 2\nGPU: 1"]
    S --> P3["๐Ÿ“ฆ Pod 3\nGPU: 1"]
    H["๐Ÿ“ˆ HPA / KEDA\nwatches queue depth"] -.->|"scale Deployment 2 โ†’ 3"| P3

    style I fill:#d8dfe8,stroke:#b0bac8
    style S fill:#e8e0d4,stroke:#c8b89a
    style H fill:#dde4dc,stroke:#b0c4b0

GPU Scheduling and Node Pools

GPU machines are expensive, so most clusters keep them in a separate pool from regular CPU machines and use "taints" to stop ordinary workloads from accidentally landing there. Your model-serving pods carry a matching "toleration" that says, in effect, "I'm allowed on the expensive machines."

GPU node pools are tainted (e.g. nvidia.com/gpu=present:NoSchedule) so only pods with a matching tolerations entry get scheduled there. The pod spec requests GPUs via resources.limits."nvidia.com/gpu" - Kubernetes' device plugin framework (the NVIDIA device plugin DaemonSet) advertises GPU capacity per node and the scheduler bin-packs accordingly. Unlike CPU/memory, GPU requests and limits must be equal - GPUs are not shareable/overcommittable by default (MIG partitioning is the exception, out of scope here).


Helm Charts

A Helm chart is a template. You write the Kubernetes YAML once with placeholders (image tag, replica count, GPU count), then supply a short values.yaml per environment. helm install fills in the placeholders and applies the result - no more maintaining four near-identical copies of the same YAML.

A chart's structure: Chart.yaml (metadata, apiVersion), values.yaml (defaults), templates/*.yaml (Go-templated Kubernetes manifests referencing .Values.x). helm template renders the manifests without applying them (useful for review/diffing in CI); helm install --dry-run validates against the live cluster's API without creating resources; helm upgrade --install is the idempotent deploy command used in most pipelines. Terraform typically owns the layer underneath the chart - VPC, the GKE/EKS/AKS cluster itself, node pools, IAM - while Helm owns what runs inside the cluster.


Study Notes

Must-know for interviews:

  • Deployment + Service + Ingress + HPA are the four objects that cover most model-serving deployments
  • GPU requests/limits must be equal - GPUs aren't overcommittable like CPU/memory
  • Taints on GPU node pools + tolerations on pods keep non-GPU workloads off expensive machines
  • Default HPA only watches CPU/memory - scaling on GPU or serving metrics needs a custom-metrics pipeline (Prometheus + adapter, or KEDA); for LLMs prefer queue depth or latency over GPU utilization
  • Helm separates "what to deploy" (templates) from "how" per environment (values.yaml) - Terraform typically owns the cluster/infra layer beneath it (Infrastructure as Code)

Check Yourself

Check yourself
0 / 5 answered
  1. A pod spec sets requests: {nvidia.com/gpu: 1} and limits: {nvidia.com/gpu: 2}. What happens?
  2. What does a NoSchedule taint on a GPU node pool do?
  3. Why can't you just set a GPU limit higher than the request, like you would for CPU?
  4. What does a taint on a GPU node pool actually do?
  5. What's the difference between helm template and helm install --dry-run?

Exercises

Exercise - Write the GPU pod spec

Write the relevant fragment of a Deployment's pod template for a vLLM server that needs one GPU, must only run on a node pool tainted nvidia.com/gpu=present:NoSchedule and labelled pool: gpu-l4, and must not be marked ready until the model has loaded (GET /health returns 200).

Solution
spec:
  nodeSelector:
    pool: gpu-l4
  tolerations:
    - key: nvidia.com/gpu
      operator: Equal
      value: present
      effect: NoSchedule
  containers:
    - name: vllm
      image: vllm/vllm-openai:<pinned-tag>
      resources:
        limits:
          nvidia.com/gpu: 1
      readinessProbe:
        httpGet: {path: /health, port: 8000}
        periodSeconds: 10
      startupProbe:
        httpGet: {path: /health, port: 8000}
        failureThreshold: 60
        periodSeconds: 10

The toleration allows the pod onto the tainted pool; the nodeSelector requires it - you need both. The startup probe gives the model up to 10 minutes to load before liveness checks could restart it.

References

Last reviewed: 2026-09

โšกAI-assisted content - always verify, always explore multiple perspectivesยท