Kubernetes and Helm
Kubernetes is the system that keeps your model-serving app running in production - it restarts crashed pods, spreads traffic across replicas, and scales up when demand spikes. Helm is the packaging layer on top: instead of hand-editing a dozen YAML files per environment, you fill in one small values file and Helm generates the rest.
A GPU inference workload on Kubernetes needs the same core objects as any web service (Deployment, Service, Ingress) plus GPU-specific scheduling concerns (node pools, resource requests, autoscaling on serving signals such as queue depth rather than CPU). Helm templates the boilerplate so the same chart deploys to dev/staging/prod with different values.yaml overrides.
- Name the four Kubernetes objects behind most model-serving deployments and what each does
- Schedule GPU pods correctly with
nvidia.com/gpulimits, node-pool taints and tolerations - Explain why the default HPA cannot scale an LLM server well and what signal to use instead
- Structure a Helm chart and choose between
helm template,--dry-runandupgrade --install
The Core Objects
Four building blocks cover most of what you need: a Deployment (how many copies of your app to run and how to update them), a Service (a stable network address for those copies), an Ingress (the public door into your cluster), and an HPA (a rule that adds more copies automatically under load).
- Deployment - declares desired replica count, the pod template (container image, resource requests/limits, GPU count via
nvidia.com/gpu: 1), and the rollout strategy (RollingUpdatewithmaxSurge/maxUnavailable) - Service - a stable ClusterIP/virtual address in front of a set of pods selected by label, so pods can restart or reschedule without clients needing to track individual pod IPs
- Ingress - routes external HTTP(S) traffic into the cluster to the right Service, usually via an ingress controller (nginx, Traefik) that terminates TLS
- HorizontalPodAutoscaler (HPA) - watches a metric (CPU/memory by default; anything else, such as GPU utilization from the DCGM exporter or vLLM's queue depth, via a custom-metrics adapter) and adjusts replica count within a min/max band. For LLM serving, queue depth is the better signal - see LLM Serving on Kubernetes
flowchart LR
U["๐ Client Request"] --> I["๐ช Ingress\nTLS termination,\nrouting"]
I --> S["๐ Service\nStable address,\nload balances pods"]
S --> P1["๐ฆ Pod 1\nGPU: 1"]
S --> P2["๐ฆ Pod 2\nGPU: 1"]
S --> P3["๐ฆ Pod 3\nGPU: 1"]
H["๐ HPA / KEDA\nwatches queue depth"] -.->|"scale Deployment 2 โ 3"| P3
style I fill:#d8dfe8,stroke:#b0bac8
style S fill:#e8e0d4,stroke:#c8b89a
style H fill:#dde4dc,stroke:#b0c4b0
GPU Scheduling and Node Pools
GPU machines are expensive, so most clusters keep them in a separate pool from regular CPU machines and use "taints" to stop ordinary workloads from accidentally landing there. Your model-serving pods carry a matching "toleration" that says, in effect, "I'm allowed on the expensive machines."
GPU node pools are tainted (e.g. nvidia.com/gpu=present:NoSchedule) so only pods with a matching tolerations entry get scheduled there. The pod spec requests GPUs via resources.limits."nvidia.com/gpu" - Kubernetes' device plugin framework (the NVIDIA device plugin DaemonSet) advertises GPU capacity per node and the scheduler bin-packs accordingly. Unlike CPU/memory, GPU requests and limits must be equal - GPUs are not shareable/overcommittable by default (MIG partitioning is the exception, out of scope here).
Helm Charts
A Helm chart is a template. You write the Kubernetes YAML once with placeholders (image tag, replica count, GPU count), then supply a short values.yaml per environment. helm install fills in the placeholders and applies the result - no more maintaining four near-identical copies of the same YAML.
A chart's structure: Chart.yaml (metadata, apiVersion), values.yaml (defaults), templates/*.yaml (Go-templated Kubernetes manifests referencing .Values.x). helm template renders the manifests without applying them (useful for review/diffing in CI); helm install --dry-run validates against the live cluster's API without creating resources; helm upgrade --install is the idempotent deploy command used in most pipelines. Terraform typically owns the layer underneath the chart - VPC, the GKE/EKS/AKS cluster itself, node pools, IAM - while Helm owns what runs inside the cluster.
Study Notes
Must-know for interviews:
- Deployment + Service + Ingress + HPA are the four objects that cover most model-serving deployments
- GPU requests/limits must be equal - GPUs aren't overcommittable like CPU/memory
- Taints on GPU node pools + tolerations on pods keep non-GPU workloads off expensive machines
- Default HPA only watches CPU/memory - scaling on GPU or serving metrics needs a custom-metrics pipeline (Prometheus + adapter, or KEDA); for LLMs prefer queue depth or latency over GPU utilization
- Helm separates "what to deploy" (templates) from "how" per environment (values.yaml) - Terraform typically owns the cluster/infra layer beneath it (Infrastructure as Code)
Check Yourself
- A pod spec sets
requests: {nvidia.com/gpu: 1}andlimits: {nvidia.com/gpu: 2}. What happens? - What does a
NoScheduletaint on a GPU node pool do? - Why can't you just set a GPU limit higher than the request, like you would for CPU?
- What does a taint on a GPU node pool actually do?
- What's the difference between
helm templateandhelm install --dry-run?
Exercises
Write the relevant fragment of a Deployment's pod template for a vLLM server that needs one GPU, must only run on a node pool tainted nvidia.com/gpu=present:NoSchedule and labelled pool: gpu-l4, and must not be marked ready until the model has loaded (GET /health returns 200).
Solution
spec:
nodeSelector:
pool: gpu-l4
tolerations:
- key: nvidia.com/gpu
operator: Equal
value: present
effect: NoSchedule
containers:
- name: vllm
image: vllm/vllm-openai:<pinned-tag>
resources:
limits:
nvidia.com/gpu: 1
readinessProbe:
httpGet: {path: /health, port: 8000}
periodSeconds: 10
startupProbe:
httpGet: {path: /health, port: 8000}
failureThreshold: 60
periodSeconds: 10
The toleration allows the pod onto the tainted pool; the nodeSelector requires it - you need both. The startup probe gives the model up to 10 minutes to load before liveness checks could restart it.
References
- Kubernetes, Schedule GPUs (2026)
- Kubernetes, Taints and Tolerations (2026)
- Kubernetes, Horizontal Pod Autoscaling (2026)
- Helm, Charts (2026)
Last reviewed: 2026-09