Contents
Map

Quiz · 02 · Prog Langs

62 questions from 12 pages

These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →

Python for AI Engineering

Check yourself
0 / 5 answered
  1. Why commit uv.lock for an application?
  2. You run 5,000 model calls with asyncio.gather and get mostly HTTP 429 errors. What's the first fix?
  3. Why add random jitter to retry backoff?
  4. Which work should go to a process pool rather than asyncio or threads in standard CPython?
  5. Why test LLM application code with fake model responses rather than real calls?

API Design for LLM Services

Check yourself
0 / 5 answered
  1. A document-analysis endpoint takes 2-6 minutes per request. Which interaction style fits best?
  2. Users report the chat answer appears all at once after 15 seconds, although the backend streams. What is the most likely cause?
  3. A client times out waiting for a POST that starts an agent run, and retries. How do you avoid running the agent twice?
  4. Which status code and header should a saturated LLM API return when it sheds load?
  5. Why should a streaming endpoint check for client disconnects?

Git Workflows for ML

Check yourself
0 / 5 answered
  1. Where should a 40 GB training dataset used by several experiments be versioned?
  2. An API key was pushed to a public repository an hour ago. What is the first thing to do?
  3. Why should prompt changes go through pull requests with an eval gate?
  4. Quality dropped somewhere in the last 64 commits. Roughly how many eval runs does git bisect need to find the culprit?
  5. Why strip notebook outputs before committing?

Linux & the GPU Box

Check yourself
0 / 5 answered
  1. nvidia-smi shows 'CUDA Version: 13.0'. What does that tell you?
  2. Your training job was killed by a spot-instance reclaim and lost six hours of progress. What change prevents that?
  3. PyTorch DataLoader workers crash with 'bus error' inside a Docker container. What is the likely fix?
  4. How do you call a vLLM server running on port 8000 of a remote GPU machine from your laptop without opening the port to the internet?
  5. dmesg shows repeated 'Xid 79' for one GPU. What does it mean, and what do you do?

Tensors & Autograd

Check yourself
0 / 8 answered
  1. You run two iterations of forward pass + loss.backward() without calling zero_grad() in between. What does w.grad contain?
  2. y = w * x + b with w=2, x=3, b=1 and loss = (y - 10)**2. What is w.grad after loss.backward()?
  3. Why do gradients accumulate by default instead of being overwritten?
  4. What happens if you call .backward() twice on the same graph without retain_graph=True?
  5. What is the difference between .detach() and torch.no_grad()?
  6. Why does model.parameters() require gradients by default but a raw input tensor doesn't?
  7. What does loss.backward() actually compute mathematically?
  8. Can you call .backward() on a non-scalar tensor?

Dataset & DataLoader

Check yourself
0 / 6 answered
  1. GPU utilisation hovers at 30% during training and a profiler shows the model waiting for batches. What do you try first?
  2. Which setting belongs on a validation DataLoader?
  3. Why does __getitem__ apply transforms instead of pre-processing the whole dataset once in __init__?
  4. What's the tradeoff of increasing num_workers?
  5. Why use drop_last=True for training but not always for validation?
  6. What does collate_fn do and when do you need a custom one?

Training Loop From Scratch

Check yourself
0 / 6 answered
  1. Which two layer types change behaviour between model.train() and model.eval()?
  2. Training loss keeps falling while validation loss rises after epoch 5. What is the most likely explanation?
  3. What's the single most common bug caused by forgetting model.eval()?
  4. Why is optimizer.zero_grad() called before backward() and not after step()?
  5. Why wrap the validation loop in torch.no_grad() even though you're not calling backward() there?
  6. How do you compute an epoch-level average loss correctly when the last batch is smaller than the others?

Checkpointing & Mixed Precision

Check yourself
0 / 6 answered
  1. Why does bfloat16 training usually not need a GradScaler?
  2. Since PyTorch 2.6, what does torch.load do by default?
  3. Why save model.state_dict() instead of torch.save(model, path)?
  4. What exactly does GradScaler do, step by step?
  5. Why does map_location=device matter when loading a checkpoint?
  6. Does mixed precision hurt model accuracy?

Debugging & GPU Memory

Check yourself
0 / 7 answered
  1. Roughly how much memory do parameters + gradients + Adam states need for a 1B-parameter model in full float32 training?
  2. Memory grows every iteration until an OOM, even at a small batch size. What is the classic cause?
  3. You get "Expected all tensors to be on the same device, but found at least two devices." What's the most common overlooked cause?
  4. What's the fastest way to isolate a shape-mismatch bug in a multi-layer model?
  5. Why does reducing batch size fix a CUDA OOM error?
  6. What does torch.cuda.empty_cache() actually do - does it free up "more" memory for training?
  7. A parameter shows param.grad is None after .backward() - what are the two most likely causes?

PyTorch for LLMs

Check yourself
0 / 3 answered
  1. Why does bf16 training not need a GradScaler while fp16 does?
  2. A 13B model's full training state doesn't fit on one 80 GB GPU. What does FSDP2 change compared with DDP?
  3. You downloaded model.pt from an unfamiliar repository. What is the risk in torch.load('model.pt', weights_only=False), and what should you do instead?

Raw PyTorch Classifier

Check yourself
0 / 3 answered
  1. The run is interrupted after epoch 3 and restarted with the checkpoint. Which state must the checkpoint contain for training to continue identically?
  2. Why is GradScaler created with enabled=use_amp in train.py?
  3. Validation accuracy is noisy between two evaluations of the same checkpoint. What is the first thing to check?

GPT From Scratch

Check yourself
0 / 3 answered
  1. Your first logged loss is 12.3 with a 256-token vocabulary. What is the most likely problem?
  2. Why does the model use 2 KV heads for 6 query heads?
  3. Why is the output projection tied to the input embedding?