Code Lab - 01: Helm Chart & Release Checklist
← Back to Overview: Production Engineering
This lab packages the Serving & Inference FastAPI + vLLM endpoint as a real deployable artifact - a Dockerfile, a Helm chart with autoscaling, and a one-page release-readiness checklist you could hand to a team on day one of an on-call rotation.
- Build a multi-stage GPU inference image that mounts weights at run time instead of copying them in
- Package the service as a Helm chart with GPU limits, taints and tolerations, probes sized for model load time, and a queue-depth HPA
- Validate the chart with
helm lint,helm templateand--dry-runbefore deploying - Turn rollout and rollback policy into a release-readiness checklist a team can follow
- Docker for GPU Inference, Kubernetes & Helm and Model Lifecycle & Rollout
- The FastAPI + vLLM Endpoint lab's
serve.py
Lab Overview
| Property | Detail |
|---|---|
| Patterns | GPU Dockerfile, Helm templating, queue-depth HPA autoscaling |
| Deploys | The Module 08 FastAPI + vLLM serving endpoint |
| Complexity | Intermediate |
| Files | Dockerfile, helm-chart/, release-readiness-checklist.md |
What's Inside
01-Helm-Chart-and-Release-Checklist/
README.mdx - architecture, what this demonstrates, how to run it
Dockerfile - multi-stage GPU inference image
helm-chart/
Chart.yaml
values.yaml - image, replica count, GPU resources, autoscaling bounds
templates/
deployment.yaml
service.yaml
hpa.yaml
release-readiness-checklist.md - the standalone "first 90 days" artifact
Prerequisites
- Docker with
nvidia-container-toolkitfor local GPU testing (optional -helm lint/--dry-rundon't need a GPU) helmv3 andkubectlpointed at any cluster (a localkind/minikubecluster is enough to validate)
Getting Started
cd 01-Helm-Chart-and-Release-Checklist
helm lint helm-chart/
helm install --dry-run vllm-serving helm-chart/ -f helm-chart/values.yaml
What to Read Alongside
- Docker for GPU Inference
- Kubernetes & Helm
- Model Lifecycle & Rollout - the checklist operationalizes this note
Check Yourself
- Why does the chart scale on vLLM's waiting-request count rather than GPU utilization?
- What does the startup probe protect against?
- Why is the model mounted from a PersistentVolumeClaim rather than copied into the image?
Exercises
Using this chart, describe how you would run a 10% canary of a new model version next to the current one, and how you would roll it back.
Hint
Two releases, one Service?
Hint
What does the Service select on?
Solution
Install a second release (helm install vllm-canary ...) with the new MODEL_PATH or image tag and 1 replica, next to 9 replicas of the stable release, and point one Service at a pod label both releases share (the chart's Service selects on the release name, so add a shared label through values.yaml) - traffic then splits roughly by replica count (about 10%). For precise percentages, use a Gateway API HTTPRoute with weighted backends or a service mesh instead. Roll back by uninstalling the canary release (or scaling it to zero); promote by upgrading the stable release to the new version and removing the canary.
Load tests (module 08 lab) show that one replica meets the SLO (p95 TTFT < 500 ms) up to 24 concurrent requests, and that above about 4 waiting requests p95 TTFT exceeds the SLO. Choose targetWaitingRequestsPerPod, minReplicas and maxReplicas for a peak of 120 concurrent requests and an overnight floor of 10.
Solution
- Target about 2-3 waiting requests per pod: below the point where TTFT breaks, leaving headroom for the minutes a new replica needs to load.
- Peak: 120 / 24 = 5 replicas at the limit; add headroom for a replica failing or loading -
maxReplicas: 7. - Floor: 10 concurrent needs 1 replica; keep
minReplicas: 2for availability during node failure or rollout. - Scale-up must lead demand because cold starts take minutes - consider pre-scaling on a schedule before known peaks.
References
- Helm, Chart best practices (2026)
- Kubernetes, Configure liveness, readiness and startup probes (2026)
- vLLM, Production metrics -
vllm:num_requests_waiting(2026) - Prometheus, Prometheus Adapter for Kubernetes metrics APIs (2026)
Last reviewed: 2026-09