Contents
Map

Quiz · 08 · Inference & Serving

49 questions from 8 pages

These are the Check Yourself questions from each page of the module, collected in course order. Each heading links back to the page the questions test. All module quizzes →

KV Cache & Inference

Check yourself
0 / 9 answered
  1. Llama 3 8B has 32 layers, 8 KV heads and a head dimension of 128. In BF16, how much KV cache does one token occupy across all layers?
  2. A model moves from multi-head attention with 32 KV heads to grouped-query attention with 8 KV heads. What happens to the KV cache?
  3. With standard speculative decoding (draft tokens verified with rejection sampling), how does the output distribution compare with running the target model alone?
  4. Why is the decode phase usually memory-bandwidth bound rather than compute bound?
  5. What two phases does LLM inference have?
  6. Why does a large batch size improve throughput but hurt latency?
  7. What is Paged Attention?
  8. How does GQA differ from MHA?
  9. When does speculative decoding NOT help?

vLLM & Paged Attention

Check yourself
0 / 7 answered
  1. What did the vLLM paper measure as KV-cache memory waste in earlier serving systems versus PagedAttention?
  2. A long prompt arrives while 50 sequences are decoding. Which setting stops its prefill from stalling their decode steps?
  3. What does vLLM's block_size control?
  4. Why is chunked prefill needed alongside continuous batching?
  5. How does prefix caching avoid recomputing a shared system prompt for every request?
  6. What's the practical effect of setting gpu_memory_utilization too low?
  7. Why does naive transformers.generate() in a loop not scale to production traffic even on a large GPU?

Quantized Inference

Check yourself
0 / 6 answered
  1. On an H100 fleet, which quantization would you try first for serving?
  2. Why does weight-only 4-bit quantization speed up decode at small batch sizes?
  3. Why can't you just serve a bitsandbytes NF4-quantized model the same way you'd serve a GPTQ/AWQ model?
  4. What does AWQ preserve that plain uniform quantization doesn't?
  5. When would INT8 be the better choice over INT4 for a production deployment?
  6. A teammate proposes reusing the QLoRA NF4 config from training directly for production serving, to save setup time. What's the issue?

Batching, Concurrency & Latency

Check yourself
0 / 6 answered
  1. Aggregate throughput keeps rising up to 128 concurrent requests, but p95 per-request tok/s falls below the SLO at 48. What is your per-GPU capacity?
  2. What two resources can each cap concurrent users per GPU?
  3. Why can concurrency exceed max_num_seqs under continuous batching?
  4. What's the risk of sizing GPU capacity purely off an aggregate throughput number?
  5. Why add more GPU replicas instead of just raising batch size indefinitely on one GPU?
  6. Your load test shows aggregate throughput still climbing at concurrency=200, but p95 per-request latency has already tripled. Do you keep raising the concurrency limit?

Streaming & TTFT

Check yourself
0 / 7 answered
  1. A RAG app sends 6,000 prompt tokens and gets 100-token answers. Which latency component dominates the time to first token?
  2. Streaming a response changes which metric?
  3. Does streaming reduce total generation time?
  4. What two components make up TTFT?
  5. Why does a long RAG context increase TTFT even if the model's actual answer is short?
  6. Why measure TTFT at increasing concurrency instead of just at concurrency=1?
  7. A user reports that responses "feel slow" even though your dashboard shows average total generation time is well within SLO. What's the likely gap in your metrics?

Production Deployment

Check yourself
0 / 7 answered
  1. A 1,000-token system prompt is shared by every request. Which cache removes most of its cost?
  2. Which metric is the better early warning of running out of serving capacity?
  3. What is TTFT and what determines it?
  4. Why is continuous batching better than static batching?
  5. What is chunked prefill?
  6. When should you use RAG vs sliding window?
  7. Name 3 ways to improve TTFT.

The Modern Serving Stack

Check yourself
0 / 4 answered
  1. A chat service must keep p99 inter-token latency under 50 ms, but long RAG prompts cause ITL spikes when their prefill runs. What is the first thing to try on a single-pool deployment?
  2. Why does KV-cache-aware routing matter most for agent workloads?
  3. Speculative decoding speeds up TPOT a lot at batch size 1 but barely at batch size 64. Why?
  4. What does goodput measure that throughput doesn't?

FastAPI + vLLM Endpoint

Check yourself
0 / 3 answered
  1. Why should the load test report p95 latency rather than the average?
  2. TTFT rises sharply at concurrency 64 while aggregate tokens/s still grows. What is happening?
  3. Why does the server use AsyncLLM rather than calling LLM.generate() per request?