Quiz · 08 · Inference & Serving
49 questions from 8 pages
These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →
KV Cache & Inference
Check yourself
0 / 9 answered
- Llama 3 8B has 32 layers, 8 KV heads and a head dimension of 128. In BF16, how much KV cache does one token occupy across all layers?
- A model moves from multi-head attention with 32 KV heads to grouped-query attention with 8 KV heads. What happens to the KV cache?
- With standard speculative decoding (draft tokens verified with rejection sampling), how does the output distribution compare with running the target model alone?
- Why is the decode phase usually memory-bandwidth bound rather than compute bound?
- What two phases does LLM inference have?
- Why does a large batch size improve throughput but hurt latency?
- What is Paged Attention?
- How does GQA differ from MHA?
- When does speculative decoding NOT help?
vLLM & Paged Attention
Check yourself
0 / 7 answered
- What did the vLLM paper measure as KV-cache memory waste in earlier serving systems versus PagedAttention?
- A long prompt arrives while 50 sequences are decoding. Which setting stops its prefill from stalling their decode steps?
- What does vLLM's
block_sizecontrol? - Why is chunked prefill needed alongside continuous batching?
- How does prefix caching avoid recomputing a shared system prompt for every request?
- What's the practical effect of setting
gpu_memory_utilizationtoo low? - Why does naive
transformers.generate()in a loop not scale to production traffic even on a large GPU?
Quantized Inference
Check yourself
0 / 6 answered
- On an H100 fleet, which quantization would you try first for serving?
- Why does weight-only 4-bit quantization speed up decode at small batch sizes?
- Why can't you just serve a
bitsandbytesNF4-quantized model the same way you'd serve a GPTQ/AWQ model? - What does AWQ preserve that plain uniform quantization doesn't?
- When would INT8 be the better choice over INT4 for a production deployment?
- A teammate proposes reusing the QLoRA NF4 config from training directly for production serving, to save setup time. What's the issue?
Batching, Concurrency & Latency
Check yourself
0 / 6 answered
- Aggregate throughput keeps rising up to 128 concurrent requests, but p95 per-request tok/s falls below the SLO at 48. What is your per-GPU capacity?
- What two resources can each cap concurrent users per GPU?
- Why can concurrency exceed
max_num_seqsunder continuous batching? - What's the risk of sizing GPU capacity purely off an aggregate throughput number?
- Why add more GPU replicas instead of just raising batch size indefinitely on one GPU?
- Your load test shows aggregate throughput still climbing at concurrency=200, but p95 per-request latency has already tripled. Do you keep raising the concurrency limit?
Streaming & TTFT
Check yourself
0 / 7 answered
- A RAG app sends 6,000 prompt tokens and gets 100-token answers. Which latency component dominates the time to first token?
- Streaming a response changes which metric?
- Does streaming reduce total generation time?
- What two components make up TTFT?
- Why does a long RAG context increase TTFT even if the model's actual answer is short?
- Why measure TTFT at increasing concurrency instead of just at concurrency=1?
- A user reports that responses "feel slow" even though your dashboard shows average total generation time is well within SLO. What's the likely gap in your metrics?
Production Deployment
Check yourself
0 / 7 answered
- A 1,000-token system prompt is shared by every request. Which cache removes most of its cost?
- Which metric is the better early warning of running out of serving capacity?
- What is TTFT and what determines it?
- Why is continuous batching better than static batching?
- What is chunked prefill?
- When should you use RAG vs sliding window?
- Name 3 ways to improve TTFT.
The Modern Serving Stack
Check yourself
0 / 4 answered
- A chat service must keep p99 inter-token latency under 50 ms, but long RAG prompts cause ITL spikes when their prefill runs. What is the first thing to try on a single-pool deployment?
- Why does KV-cache-aware routing matter most for agent workloads?
- Speculative decoding speeds up TPOT a lot at batch size 1 but barely at batch size 64. Why?
- What does goodput measure that throughput doesn't?
FastAPI + vLLM Endpoint
Check yourself
0 / 3 answered
- Why should the load test report p95 latency rather than the average?
- TTFT rises sharply at concurrency 64 while aggregate tokens/s still grows. What is happening?
- Why does the server use AsyncLLM rather than calling LLM.generate() per request?