Quiz · 09 · Production Engineering
47 questions from 10 pages
These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
Docker for GPU Inference
Check yourself
0 / 6 answered
- Where does
libcuda.so(the GPU driver library) come from inside a running GPU container? - Which situation genuinely requires a CUDA
-develimage (in at least one build stage)? - Why does
torch.cuda.is_available()returnFalseinside a container even on a GPU host? - What's the practical difference between a CUDA
-develand-runtimeimage tag? - Why should model weights usually not be
COPY'd into the Docker image? - A teammate suggests
FROM python:3.12-slimand installing CUDAtorchvia pip, arguing it's simpler than a CUDA base image. Is that wrong?
Kubernetes & Helm
Check yourself
0 / 5 answered
- A pod spec sets
requests: {nvidia.com/gpu: 1}andlimits: {nvidia.com/gpu: 2}. What happens? - What does a
NoScheduletaint on a GPU node pool do? - Why can't you just set a GPU limit higher than the request, like you would for CPU?
- What does a taint on a GPU node pool actually do?
- What's the difference between
helm templateandhelm install --dry-run?
Model Lifecycle & Rollout
Check yourself
0 / 5 answered
- You want to see how a new model behaves on real traffic without any user seeing its answers. Which strategy fits?
- In MLflow's current registry, how is "the version production should serve" usually marked?
- What's the difference between canary and champion-challenger?
- Why is "we'll decide if it's bad enough to roll back when we see it" a risky plan?
- What does a model registry actually store?
Security & Compliance
Check yourself
0 / 6 answered
- Why is loading a
.bin/.ptcheckpoint from an untrusted hub repository risky? - A RAG assistant returns passages from HR documents the user is not permitted to read. Which OWASP LLM risk is this, and what is the first-line control?
- Why is a shared "AI service account" with broad permissions risky?
- Why do tracing tools pose a specific compliance risk for LLM systems that they don't for typical apps?
- What are HIPAA's five technical safeguards?
- What should trigger a CI pipeline to block a merge, from a security standpoint?
LLM Serving on Kubernetes
Check yourself
0 / 3 answered
- A vLLM deployment shows 98% GPU utilization at both 3 AM and peak hour, but TTFT at peak is 10× worse. What should autoscaling watch instead?
- You need to serve 12 small embedding and reranking models with strict isolation on a few H100s. Which GPU-sharing approach fits best?
- What problem does LeaderWorkerSet solve that a Deployment doesn't?
Observability, SLOs & Incidents
Check yourself
0 / 5 answered
- A chat service has a 99.5% goodput SLO over 30 days. Using the standard fast-burn page (2% of the budget in 1 hour), above what error rate does the page fire?
- Why does a multi-window burn-rate alert require both a long and a short window to exceed the threshold?
- Every dashboard is green - no errors, latency within SLO - but user thumbs-down rate has doubled since Tuesday. What kind of incident is this, and what is the first thing to check?
- Why is GPU utilization or queue depth a poor SLI but a good early-warning alert?
- Which of these changes should go through the error-budget policy and a canary like a code deploy?
Cloud Networking & IAM for AI
Check yourself
0 / 5 answered
- Your RAG API calls Amazon Bedrock. Which setup keeps traffic off the public internet and makes a leaked credential useless from outside your network?
- Why prefer workload identity (EKS Pod Identity, GKE Workload Identity Federation, Entra Workload ID) over access keys?
- An agent with a web-browsing tool runs in a private subnet. How should its outbound access be designed?
- What does a data perimeter (e.g. VPC Service Controls) protect against that least-privilege IAM alone does not?
- Why should model weights be pre-staged in your own storage rather than downloaded from a public hub at node start-up?
Infrastructure as Code
Check yourself
0 / 5 answered
- A reviewer sees a one-line HCL change to a GPU node pool's machine_type. What should they check before approving?
- Why must Terraform/OpenTofu state be stored remotely with locking?
- A scheduled
plan -detailed-exitcodeexits with code 2 overnight. What does it mean and what do you do? - Why is the state file security-sensitive?
- Which tool should deploy a new version of the vLLM server every few days: IaC or GitOps?
Helm Chart & Release Checklist
Check yourself
0 / 3 answered
- Why does the chart scale on vLLM's waiting-request count rather than GPU utilization?
- What does the startup probe protect against?
- Why is the model mounted from a PersistentVolumeClaim rather than copied into the image?
Terraform Private Endpoint & SLO Alerts
Check yourself
0 / 4 answered
- The role policy requires aws:SourceVpce to equal the endpoint id. What does that protect against?
- Why does the lab test that a healthy 0.2% error rate produces no alerts?
- With a 99.5% availability SLO, what error rate triggers the fast-burn page, and why does it need both the 1 h and 5 m windows?
- The first tofu test run failed with 'Invalid ARN Value'. Was that a bug in main.tf?