Quiz · 02 · Prog Langs
62 questions from 12 pages
These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
Python for AI Engineering
Check yourself
0 / 5 answered
- Why commit uv.lock for an application?
- You run 5,000 model calls with asyncio.gather and get mostly HTTP 429 errors. What's the first fix?
- Why add random jitter to retry backoff?
- Which work should go to a process pool rather than asyncio or threads in standard CPython?
- Why test LLM application code with fake model responses rather than real calls?
API Design for LLM Services
Check yourself
0 / 5 answered
- A document-analysis endpoint takes 2-6 minutes per request. Which interaction style fits best?
- Users report the chat answer appears all at once after 15 seconds, although the backend streams. What is the most likely cause?
- A client times out waiting for a POST that starts an agent run, and retries. How do you avoid running the agent twice?
- Which status code and header should a saturated LLM API return when it sheds load?
- Why should a streaming endpoint check for client disconnects?
Git Workflows for ML
Check yourself
0 / 5 answered
- Where should a 40 GB training dataset used by several experiments be versioned?
- An API key was pushed to a public repository an hour ago. What is the first thing to do?
- Why should prompt changes go through pull requests with an eval gate?
- Quality dropped somewhere in the last 64 commits. Roughly how many eval runs does git bisect need to find the culprit?
- Why strip notebook outputs before committing?
Linux & the GPU Box
Check yourself
0 / 5 answered
- nvidia-smi shows 'CUDA Version: 13.0'. What does that tell you?
- Your training job was killed by a spot-instance reclaim and lost six hours of progress. What change prevents that?
- PyTorch DataLoader workers crash with 'bus error' inside a Docker container. What is the likely fix?
- How do you call a vLLM server running on port 8000 of a remote GPU machine from your laptop without opening the port to the internet?
- dmesg shows repeated 'Xid 79' for one GPU. What does it mean, and what do you do?
Tensors & Autograd
Check yourself
0 / 8 answered
- You run two iterations of forward pass +
loss.backward()without callingzero_grad()in between. What doesw.gradcontain? y = w * x + bwith w=2, x=3, b=1 andloss = (y - 10)**2. What isw.gradafterloss.backward()?- Why do gradients accumulate by default instead of being overwritten?
- What happens if you call
.backward()twice on the same graph withoutretain_graph=True? - What is the difference between
.detach()andtorch.no_grad()? - Why does
model.parameters()require gradients by default but a raw input tensor doesn't? - What does
loss.backward()actually compute mathematically? - Can you call
.backward()on a non-scalar tensor?
Dataset & DataLoader
Check yourself
0 / 6 answered
- GPU utilisation hovers at 30% during training and a profiler shows the model waiting for batches. What do you try first?
- Which setting belongs on a validation DataLoader?
- Why does
__getitem__apply transforms instead of pre-processing the whole dataset once in__init__? - What's the tradeoff of increasing
num_workers? - Why use
drop_last=Truefor training but not always for validation? - What does
collate_fndo and when do you need a custom one?
Training Loop From Scratch
Check yourself
0 / 6 answered
- Which two layer types change behaviour between model.train() and model.eval()?
- Training loss keeps falling while validation loss rises after epoch 5. What is the most likely explanation?
- What's the single most common bug caused by forgetting
model.eval()? - Why is
optimizer.zero_grad()called beforebackward()and not afterstep()? - Why wrap the validation loop in
torch.no_grad()even though you're not callingbackward()there? - How do you compute an epoch-level average loss correctly when the last batch is smaller than the others?
Checkpointing & Mixed Precision
Check yourself
0 / 6 answered
- Why does bfloat16 training usually not need a GradScaler?
- Since PyTorch 2.6, what does torch.load do by default?
- Why save
model.state_dict()instead oftorch.save(model, path)? - What exactly does
GradScalerdo, step by step? - Why does
map_location=devicematter when loading a checkpoint? - Does mixed precision hurt model accuracy?
Debugging & GPU Memory
Check yourself
0 / 7 answered
- Roughly how much memory do parameters + gradients + Adam states need for a 1B-parameter model in full float32 training?
- Memory grows every iteration until an OOM, even at a small batch size. What is the classic cause?
- You get "Expected all tensors to be on the same device, but found at least two devices." What's the most common overlooked cause?
- What's the fastest way to isolate a shape-mismatch bug in a multi-layer model?
- Why does reducing batch size fix a CUDA OOM error?
- What does
torch.cuda.empty_cache()actually do - does it free up "more" memory for training? - A parameter shows
param.grad is Noneafter.backward()- what are the two most likely causes?
PyTorch for LLMs
Check yourself
0 / 3 answered
- Why does bf16 training not need a GradScaler while fp16 does?
- A 13B model's full training state doesn't fit on one 80 GB GPU. What does FSDP2 change compared with DDP?
- You downloaded
model.ptfrom an unfamiliar repository. What is the risk intorch.load('model.pt', weights_only=False), and what should you do instead?
Raw PyTorch Classifier
Check yourself
0 / 3 answered
- The run is interrupted after epoch 3 and restarted with the checkpoint. Which state must the checkpoint contain for training to continue identically?
- Why is GradScaler created with
enabled=use_ampin train.py? - Validation accuracy is noisy between two evaluations of the same checkpoint. What is the first thing to check?
GPT From Scratch
Check yourself
0 / 3 answered
- Your first logged loss is 12.3 with a 256-token vocabulary. What is the most likely problem?
- Why does the model use 2 KV heads for 6 query heads?
- Why is the output projection tied to the input embedding?