Appendix - Production Engineering
What We Learned
- Package GPU services with compatible drivers/runtimes; schedule and scale them on Kubernetes.
- Version models and roll out gradually with predefined rollback conditions.
- Protect identities, secrets, sensitive data, dependencies and network boundaries.
- Use traces, metrics and SLO burn-rate alerts to detect and manage incidents.
- Manage infrastructure through reviewed plans, locked state, policy and drift checks.
Key Acronyms, Concepts & Jargon
| Term | Short meaning |
|---|---|
| Docker / Kubernetes / Helm | Container packaging / workload orchestration / templated Kubernetes releases. |
| HPA / GPU | Horizontal Pod Autoscaler / Graphics Processing Unit: replica scaling / accelerator resource. |
| Canary / shadow rollout | Limited live release / test new behavior alongside the live system. |
| IAM / RBAC | Identity and Access Management / Role-Based Access Control: identity permissions / permissions by role. |
| PII / PHI | Personally Identifiable Information / Protected Health Information: sensitive identity / health data. |
| VPC / private endpoint | Virtual Private Cloud / private network access to a service. |
| Workload identity | Short-lived service identity instead of embedded long-lived keys. |
| IaC / GitOps | Infrastructure as Code / reconciling deployed configuration from Git. |
| Terraform state / drift | Record of managed resources / mismatch between declared and actual resources. |
| SLI / SLO / SLA | Service Level Indicator / Objective / Agreement: measurement / target / commitment. |
| Error budget / burn rate | Allowed failures / speed at which that allowance is consumed. |
| OTel | OpenTelemetry: standard instrumentation for traces, metrics and logs. |
| Runbook / postmortem | Incident procedure / review of causes, impact and corrective actions. |
| SBOM | Software Bill of Materials: inventory of software components and dependencies. |