Contents
Map

09 · Production Engineering

Appendix - Summary & Key Terms

View as:

Appendix - Production Engineering

What We Learned

  • Package GPU services with compatible drivers/runtimes; schedule and scale them on Kubernetes.
  • Version models and roll out gradually with predefined rollback conditions.
  • Protect identities, secrets, sensitive data, dependencies and network boundaries.
  • Use traces, metrics and SLO burn-rate alerts to detect and manage incidents.
  • Manage infrastructure through reviewed plans, locked state, policy and drift checks.

Key Acronyms, Concepts & Jargon

TermShort meaning
Docker / Kubernetes / HelmContainer packaging / workload orchestration / templated Kubernetes releases.
HPA / GPUHorizontal Pod Autoscaler / Graphics Processing Unit: replica scaling / accelerator resource.
Canary / shadow rolloutLimited live release / test new behavior alongside the live system.
IAM / RBACIdentity and Access Management / Role-Based Access Control: identity permissions / permissions by role.
PII / PHIPersonally Identifiable Information / Protected Health Information: sensitive identity / health data.
VPC / private endpointVirtual Private Cloud / private network access to a service.
Workload identityShort-lived service identity instead of embedded long-lived keys.
IaC / GitOpsInfrastructure as Code / reconciling deployed configuration from Git.
Terraform state / driftRecord of managed resources / mismatch between declared and actual resources.
SLI / SLO / SLAService Level Indicator / Objective / Agreement: measurement / target / commitment.
Error budget / burn rateAllowed failures / speed at which that allowance is consumed.
OTelOpenTelemetry: standard instrumentation for traces, metrics and logs.
Runbook / postmortemIncident procedure / review of causes, impact and corrective actions.
SBOMSoftware Bill of Materials: inventory of software components and dependencies.

Back to section overview

⚡AI-assisted content - always verify, always explore multiple perspectives·