Contents
Map

Quiz · 04 · Pretraining at Scale

36 questions from 9 pages

These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →

Pretraining Overview

Check yourself
0 / 4 answered
  1. Why did Llama 3 move from a 32K SentencePiece vocabulary to a 128K tiktoken-based BPE vocabulary?
  2. In Hugging Face causal-LM training, what should labels be set to at padding positions?
  3. In BERT's masked-LM objective, what happens to the 15% of selected tokens?
  4. A 1B-parameter model is trained on 20B tokens. Roughly how many training FLOPs is that?

GPU Memory & Hardware

Check yourself
0 / 4 answered
  1. How much memory does model state need to train a 13B model with mixed-precision Adam, before activations?
  2. With ZeRO stage 3 across 8 GPUs, how many bytes of model state per parameter does each GPU hold?
  3. Why is tensor parallelism usually kept within one NVLink-connected node?
  4. A 70B model is quantized to 4-bit weights for serving on one 80 GB GPU. What still limits how many users it can serve?

Data Curation & Mixtures

Check yourself
0 / 3 answered
  1. With MinHash LSH using b bands of r rows, what does increasing r (with b fixed) do?
  2. Why can annealing a small model on a candidate dataset estimate that dataset's value cheaply?
  3. Give two risks of aggressive model-based quality filtering.

Scaling Laws & Compute Budgets

Check yourself
0 / 4 answered
  1. A dense 7B model is trained on 2T tokens. Roughly how many training FLOPs is that?
  2. Why is Llama 3 8B trained on ~1,900 tokens per parameter instead of Chinchilla's ~20?
  3. For an MoE model with 671B total and 37B active parameters trained on 14.8T tokens, which N goes into C ≈ 6ND?
  4. Why did Kaplan et al. and Chinchilla reach different conclusions about the optimal parameter/data split?

Distributed Training at Scale

Check yourself
0 / 4 answered
  1. Why is tensor parallelism normally kept within a single NVLink domain?
  2. A run has 16 pipeline stages and 16 micro-batches per step with a plain 1F1B schedule. Roughly what fraction of time is the pipeline bubble?
  3. Why did DeepSeek-V3 use fine-grained (tile and block) FP8 scaling instead of one scale per tensor?
  4. A cluster of 1,024 H100s trains a 70B dense model at 1.0M tokens/second. What is the MFU in BF16?

Training Stability & Optimizers

Check yourself
0 / 4 answered
  1. Why does a warmup-stable-decay (WSD) schedule make continual pretraining easier than cosine?
  2. Attention logits keep growing during a large run and training becomes unstable. Which technique targets that directly?
  3. What problem does μP solve?
  4. What does Muon do differently from AdamW for a hidden-layer weight matrix?

Accelerators & Interconnects

Check yourself
0 / 4 answered
  1. A spec sheet lists 3,958 TFLOPS FP8 for a GPU 'with sparsity'. What figure should you use to estimate dense LLM training throughput?
  2. Moving an inference service from H100 to H200 (same compute, more memory bandwidth) mostly speeds up which phase?
  3. Why does a 72-GPU NVLink domain (GB200 NVL72) matter for serving large MoE models?
  4. What does NVFP4 do differently from MXFP4 to lose less accuracy?

CUDA Concepts & GPU Profiling

Check yourself
0 / 5 answered
  1. How many threads are in a CUDA warp, and what do they share?
  2. You time a training step with time.time() before and after the GPU calls and get 2 ms, but throughput suggests 40 ms per step. What happened?
  3. Why does fusing three elementwise operations into one kernel speed them up, when the arithmetic is identical?
  4. Nsight Systems shows the GPU idle for 30% of each step, with gaps between kernels and the main CPU thread busy in the data loader. What do you do?
  5. What is the arithmetic intensity of batch-1 decode for a 4096 x 4096 BF16 weight matrix, and what does it imply?

Profile and Fuse

Check yourself
0 / 4 answered
  1. In Part A, matrix multiplies take about half the CPU time, yet the lab targets the elementwise tail. Why?
  2. Why does the benchmark interleave the variants trial by trial rather than running all eager trials, then all compiled trials?
  3. torch.compile's output differs from eager by 3.8e-6. Is that a bug?
  4. Why does the Triton kernel use one program per row instead of splitting each row into many blocks?