Contents
Map

09 · Production Engineering

Overview

View as:

09 - Production Engineering

Taking a serving stack to production: GPU container images, Kubernetes and Helm for GPU workloads, model registries and safe rollouts, security and compliance for AI systems, the Kubernetes-native tooling for serving LLMs at scale, and operating LLM services with SLOs and incident response.

Learning objectives 9-11 hours (notes + lab)
By the end of this module you will be able to:
  • Build GPU inference images and explain where the driver, CUDA libraries and framework each come from
  • Deploy a model server on Kubernetes with correct GPU scheduling, probes and queue-depth autoscaling, packaged as a Helm chart
  • Plan a rollout with a registry, a strategy (blue/green, canary, shadow, champion-challenger) and quantified rollback criteria
  • Secure an AI system - IAM, secrets, PHI/PII redaction, OWASP LLM Top 10, model supply chain - and place it on the EU AI Act timeline
  • Choose Kubernetes-native LLM serving tools (GPU sharing and DRA, KServe, LeaderWorkerSet, llm-d, inference gateways) for a workload
  • Define SLIs and SLOs for an LLM service, alert on error-budget burn, and run incidents and blameless postmortems
  • Lay out private networking, workload identity and data perimeters for AI workloads, and manage them as reviewed, policy-checked infrastructure as code

What You Will Learn

  • Dockerfiles for GPU inference: where the driver and CUDA libraries come from, -devel vs -runtime, multi-stage builds, why weights shouldn't be baked into the image
  • Core Kubernetes objects for model serving: Deployment, Service, Ingress, HorizontalPodAutoscaler
  • GPU scheduling: node pools, taints/tolerations, why GPU requests must equal limits
  • Helm charts: templating deployments across environments, helm template vs --dry-run
  • Model lifecycle: registries (MLflow), blue/green vs canary vs champion-challenger rollout, rollback criteria decided before launch
  • Security and compliance for AI systems: IAM least privilege, secrets management, PHI/PII in traces and eval datasets, SAST/dependency scanning, HIPAA's five technical safeguards
  • Building a release-readiness checklist you'd introduce in a first 90 days
  • Kubernetes for LLMs: GPU sharing and DRA, multi-node replicas, model-aware routing, queue-depth autoscaling, cold starts
  • LLM-specific security (OWASP LLM Top 10), model supply-chain hygiene, and the EU AI Act timeline
  • Operating LLM services: SLIs for latency, goodput, quality, safety and cost; error budgets and burn-rate alerts; incident response and postmortems
  • Cloud networking and IAM for AI: private endpoints, egress control, workload identity, data perimeters, landing zones
  • Infrastructure as code with Terraform/OpenTofu: state, modules, GPU node pools, CI plans, drift and policy as code

Chapter Map

#FileTopicDifficulty
1Docker for GPU InferenceDriver vs CUDA libraries, base-image choice, multi-stage builds, weight mountingIntermediate
2Kubernetes & HelmDeployment/Service/Ingress/HPA, GPU scheduling, Helm chartsAdvanced
3Model Lifecycle & RolloutRegistries, blue/green, canary, champion-challenger, rollback criteriaAdvanced
4Security & ComplianceIAM, secrets, PHI/PII handling, SAST, HIPAA safeguards, OWASP LLM Top 10, model supply chain, EU AI Act timelineAdvanced
5LLM Serving on KubernetesGPU Operator, MIG / time-slicing / DRA, KServe, LeaderWorkerSet, Gateway API Inference Extension, llm-d, queue-depth autoscaling, cold startsAdvanced
6LLM Observability, SLOs & Incident ResponseSLIs for LLM services, SLOs and error budgets, burn-rate alerting, OpenTelemetry GenAI signals, LLM-specific incidents, incident roles and runbooks, blameless postmortemsAdvanced
7Cloud Networking & IAM for AIVPCs and subnets, private endpoints to Bedrock / Vertex AI / Azure OpenAI, egress control, workload identity, data perimeters, landing zonesAdvanced
8Infrastructure as CodeTerraform vs OpenTofu, plan/apply, remote state and locking, GPU node pools as code, modules and environments, drift, policy as code, GitOps boundaryAdvanced
9Q&A Review Bank33 Q&A pairs across all topicsAll levels

Path A: Deployment Fundamentals

  1. Docker for GPU Inference - package the service
  2. Kubernetes & Helm - run it in production
  3. Helm Chart & Release Checklist - build the real artifact
  4. Terraform Private Endpoint & SLO Alerts - private access and error-budget alerts, tested offline

Path B: Interview Preparation (Accelerated)

  1. Kubernetes & Helm - GPU scheduling questions are common at Director/Principal level
  2. Model Lifecycle & Rollout - canary vs champion-challenger is a frequent distinction to defend
  3. Security & Compliance - HIPAA safeguards named precisely
  4. LLM Observability, SLOs & Incident Response - error budgets and burn-rate math come up in SRE-flavoured rounds
  5. Q&A Review Bank - drill all questions

Path C: Regulated/Healthcare Roles

  1. Security & Compliance
  2. Model Lifecycle & Rollout - rollback discipline matters more under regulatory review
  3. Cross-reference Prior Authorization for a worked PHI-handling example

Resources

Key Cross-References

Section Appendix

Summary & Key Terms - a quick recap of this section and its essential vocabulary.


Next Topic

Previous: 08 - Inference & Serving · Next: 10 - Cloud Platforms

⚡AI-assisted content - always verify, always explore multiple perspectives·