Contents
Map

09 · Production Engineering

Helm Chart & Release Checklist

View as:

Code Lab - 01: Helm Chart & Release Checklist

← Back to Overview: Production Engineering

This lab packages the Serving & Inference FastAPI + vLLM endpoint as a real deployable artifact - a Dockerfile, a Helm chart with autoscaling, and a one-page release-readiness checklist you could hand to a team on day one of an on-call rotation.

Learning objectives 2 hours
By the end of this page you will be able to:
  • Build a multi-stage GPU inference image that mounts weights at run time instead of copying them in
  • Package the service as a Helm chart with GPU limits, taints and tolerations, probes sized for model load time, and a queue-depth HPA
  • Validate the chart with helm lint, helm template and --dry-run before deploying
  • Turn rollout and rollback policy into a release-readiness checklist a team can follow

Lab Overview

PropertyDetail
PatternsGPU Dockerfile, Helm templating, queue-depth HPA autoscaling
DeploysThe Module 08 FastAPI + vLLM serving endpoint
ComplexityIntermediate
FilesDockerfile, helm-chart/, release-readiness-checklist.md

What's Inside

01-Helm-Chart-and-Release-Checklist/
  README.mdx                        - architecture, what this demonstrates, how to run it
  Dockerfile                        - multi-stage GPU inference image
  helm-chart/
    Chart.yaml
    values.yaml                     - image, replica count, GPU resources, autoscaling bounds
    templates/
      deployment.yaml
      service.yaml
      hpa.yaml
  release-readiness-checklist.md    - the standalone "first 90 days" artifact

Prerequisites

  • Docker with nvidia-container-toolkit for local GPU testing (optional - helm lint/--dry-run don't need a GPU)
  • helm v3 and kubectl pointed at any cluster (a local kind/minikube cluster is enough to validate)

Getting Started

cd 01-Helm-Chart-and-Release-Checklist
helm lint helm-chart/
helm install --dry-run vllm-serving helm-chart/ -f helm-chart/values.yaml

What to Read Alongside

Check Yourself

Check yourself
0 / 3 answered
  1. Why does the chart scale on vLLM's waiting-request count rather than GPU utilization?
  2. What does the startup probe protect against?
  3. Why is the model mounted from a PersistentVolumeClaim rather than copied into the image?

Exercises

Exercise - Deploy a canary with the chart

Using this chart, describe how you would run a 10% canary of a new model version next to the current one, and how you would roll it back.

Hint

Two releases, one Service?

Hint

What does the Service select on?

Solution

Install a second release (helm install vllm-canary ...) with the new MODEL_PATH or image tag and 1 replica, next to 9 replicas of the stable release, and point one Service at a pod label both releases share (the chart's Service selects on the release name, so add a shared label through values.yaml) - traffic then splits roughly by replica count (about 10%). For precise percentages, use a Gateway API HTTPRoute with weighted backends or a service mesh instead. Roll back by uninstalling the canary release (or scaling it to zero); promote by upgrading the stable release to the new version and removing the canary.

Exercise - Size the HPA target

Load tests (module 08 lab) show that one replica meets the SLO (p95 TTFT < 500 ms) up to 24 concurrent requests, and that above about 4 waiting requests p95 TTFT exceeds the SLO. Choose targetWaitingRequestsPerPod, minReplicas and maxReplicas for a peak of 120 concurrent requests and an overnight floor of 10.

Solution
  • Target about 2-3 waiting requests per pod: below the point where TTFT breaks, leaving headroom for the minutes a new replica needs to load.
  • Peak: 120 / 24 = 5 replicas at the limit; add headroom for a replica failing or loading - maxReplicas: 7.
  • Floor: 10 concurrent needs 1 replica; keep minReplicas: 2 for availability during node failure or rollout.
  • Scale-up must lead demand because cold starts take minutes - consider pre-scaling on a schedule before known peaks.

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·