Contents
Map

Quiz · 09 · Production Engineering

47 questions from 10 pages

These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →

Docker for GPU Inference

Check yourself
0 / 6 answered
  1. Where does libcuda.so (the GPU driver library) come from inside a running GPU container?
  2. Which situation genuinely requires a CUDA -devel image (in at least one build stage)?
  3. Why does torch.cuda.is_available() return False inside a container even on a GPU host?
  4. What's the practical difference between a CUDA -devel and -runtime image tag?
  5. Why should model weights usually not be COPY'd into the Docker image?
  6. A teammate suggests FROM python:3.12-slim and installing CUDA torch via pip, arguing it's simpler than a CUDA base image. Is that wrong?

Kubernetes & Helm

Check yourself
0 / 5 answered
  1. A pod spec sets requests: {nvidia.com/gpu: 1} and limits: {nvidia.com/gpu: 2}. What happens?
  2. What does a NoSchedule taint on a GPU node pool do?
  3. Why can't you just set a GPU limit higher than the request, like you would for CPU?
  4. What does a taint on a GPU node pool actually do?
  5. What's the difference between helm template and helm install --dry-run?

Model Lifecycle & Rollout

Check yourself
0 / 5 answered
  1. You want to see how a new model behaves on real traffic without any user seeing its answers. Which strategy fits?
  2. In MLflow's current registry, how is "the version production should serve" usually marked?
  3. What's the difference between canary and champion-challenger?
  4. Why is "we'll decide if it's bad enough to roll back when we see it" a risky plan?
  5. What does a model registry actually store?

Security & Compliance

Check yourself
0 / 6 answered
  1. Why is loading a .bin/.pt checkpoint from an untrusted hub repository risky?
  2. A RAG assistant returns passages from HR documents the user is not permitted to read. Which OWASP LLM risk is this, and what is the first-line control?
  3. Why is a shared "AI service account" with broad permissions risky?
  4. Why do tracing tools pose a specific compliance risk for LLM systems that they don't for typical apps?
  5. What are HIPAA's five technical safeguards?
  6. What should trigger a CI pipeline to block a merge, from a security standpoint?

LLM Serving on Kubernetes

Check yourself
0 / 3 answered
  1. A vLLM deployment shows 98% GPU utilization at both 3 AM and peak hour, but TTFT at peak is 10× worse. What should autoscaling watch instead?
  2. You need to serve 12 small embedding and reranking models with strict isolation on a few H100s. Which GPU-sharing approach fits best?
  3. What problem does LeaderWorkerSet solve that a Deployment doesn't?

Observability, SLOs & Incidents

Check yourself
0 / 5 answered
  1. A chat service has a 99.5% goodput SLO over 30 days. Using the standard fast-burn page (2% of the budget in 1 hour), above what error rate does the page fire?
  2. Why does a multi-window burn-rate alert require both a long and a short window to exceed the threshold?
  3. Every dashboard is green - no errors, latency within SLO - but user thumbs-down rate has doubled since Tuesday. What kind of incident is this, and what is the first thing to check?
  4. Why is GPU utilization or queue depth a poor SLI but a good early-warning alert?
  5. Which of these changes should go through the error-budget policy and a canary like a code deploy?

Cloud Networking & IAM for AI

Check yourself
0 / 5 answered
  1. Your RAG API calls Amazon Bedrock. Which setup keeps traffic off the public internet and makes a leaked credential useless from outside your network?
  2. Why prefer workload identity (EKS Pod Identity, GKE Workload Identity Federation, Entra Workload ID) over access keys?
  3. An agent with a web-browsing tool runs in a private subnet. How should its outbound access be designed?
  4. What does a data perimeter (e.g. VPC Service Controls) protect against that least-privilege IAM alone does not?
  5. Why should model weights be pre-staged in your own storage rather than downloaded from a public hub at node start-up?

Infrastructure as Code

Check yourself
0 / 5 answered
  1. A reviewer sees a one-line HCL change to a GPU node pool's machine_type. What should they check before approving?
  2. Why must Terraform/OpenTofu state be stored remotely with locking?
  3. A scheduled plan -detailed-exitcode exits with code 2 overnight. What does it mean and what do you do?
  4. Why is the state file security-sensitive?
  5. Which tool should deploy a new version of the vLLM server every few days: IaC or GitOps?

Helm Chart & Release Checklist

Check yourself
0 / 3 answered
  1. Why does the chart scale on vLLM's waiting-request count rather than GPU utilization?
  2. What does the startup probe protect against?
  3. Why is the model mounted from a PersistentVolumeClaim rather than copied into the image?

Terraform Private Endpoint & SLO Alerts

Check yourself
0 / 4 answered
  1. The role policy requires aws:SourceVpce to equal the endpoint id. What does that protect against?
  2. Why does the lab test that a healthy 0.2% error rate produces no alerts?
  3. With a 99.5% availability SLO, what error rate triggers the fast-burn page, and why does it need both the 1 h and 5 m windows?
  4. The first tofu test run failed with 'Invalid ARN Value'. Was that a bug in main.tf?