Quiz · 04 · Pretraining at Scale
36 questions from 9 pages
These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
Pretraining Overview
Check yourself
0 / 4 answered
- Why did Llama 3 move from a 32K SentencePiece vocabulary to a 128K tiktoken-based BPE vocabulary?
- In Hugging Face causal-LM training, what should
labelsbe set to at padding positions? - In BERT's masked-LM objective, what happens to the 15% of selected tokens?
- A 1B-parameter model is trained on 20B tokens. Roughly how many training FLOPs is that?
GPU Memory & Hardware
Check yourself
0 / 4 answered
- How much memory does model state need to train a 13B model with mixed-precision Adam, before activations?
- With ZeRO stage 3 across 8 GPUs, how many bytes of model state per parameter does each GPU hold?
- Why is tensor parallelism usually kept within one NVLink-connected node?
- A 70B model is quantized to 4-bit weights for serving on one 80 GB GPU. What still limits how many users it can serve?
Data Curation & Mixtures
Check yourself
0 / 3 answered
- With MinHash LSH using b bands of r rows, what does increasing r (with b fixed) do?
- Why can annealing a small model on a candidate dataset estimate that dataset's value cheaply?
- Give two risks of aggressive model-based quality filtering.
Scaling Laws & Compute Budgets
Check yourself
0 / 4 answered
- A dense 7B model is trained on 2T tokens. Roughly how many training FLOPs is that?
- Why is Llama 3 8B trained on ~1,900 tokens per parameter instead of Chinchilla's ~20?
- For an MoE model with 671B total and 37B active parameters trained on 14.8T tokens, which N goes into C ≈ 6ND?
- Why did Kaplan et al. and Chinchilla reach different conclusions about the optimal parameter/data split?
Distributed Training at Scale
Check yourself
0 / 4 answered
- Why is tensor parallelism normally kept within a single NVLink domain?
- A run has 16 pipeline stages and 16 micro-batches per step with a plain 1F1B schedule. Roughly what fraction of time is the pipeline bubble?
- Why did DeepSeek-V3 use fine-grained (tile and block) FP8 scaling instead of one scale per tensor?
- A cluster of 1,024 H100s trains a 70B dense model at 1.0M tokens/second. What is the MFU in BF16?
Training Stability & Optimizers
Check yourself
0 / 4 answered
- Why does a warmup-stable-decay (WSD) schedule make continual pretraining easier than cosine?
- Attention logits keep growing during a large run and training becomes unstable. Which technique targets that directly?
- What problem does μP solve?
- What does Muon do differently from AdamW for a hidden-layer weight matrix?
Accelerators & Interconnects
Check yourself
0 / 4 answered
- A spec sheet lists 3,958 TFLOPS FP8 for a GPU 'with sparsity'. What figure should you use to estimate dense LLM training throughput?
- Moving an inference service from H100 to H200 (same compute, more memory bandwidth) mostly speeds up which phase?
- Why does a 72-GPU NVLink domain (GB200 NVL72) matter for serving large MoE models?
- What does NVFP4 do differently from MXFP4 to lose less accuracy?
CUDA Concepts & GPU Profiling
Check yourself
0 / 5 answered
- How many threads are in a CUDA warp, and what do they share?
- You time a training step with time.time() before and after the GPU calls and get 2 ms, but throughput suggests 40 ms per step. What happened?
- Why does fusing three elementwise operations into one kernel speed them up, when the arithmetic is identical?
- Nsight Systems shows the GPU idle for 30% of each step, with gaps between kernels and the main CPU thread busy in the data loader. What do you do?
- What is the arithmetic intensity of batch-1 decode for a 4096 x 4096 BF16 weight matrix, and what does it imply?
Profile and Fuse
Check yourself
0 / 4 answered
- In Part A, matrix multiplies take about half the CPU time, yet the lab targets the elementwise tail. Why?
- Why does the benchmark interleave the variants trial by trial rather than running all eager trials, then all compiled trials?
- torch.compile's output differs from eager by 3.8e-6. Is that a bug?
- Why does the Triton kernel use one program per row instead of splitting each row into many blocks?