Contents
Map

08 · Inference & Serving

Appendix - Summary & Key Terms

View as:

Appendix - Inference & Serving

What We Learned

  • Prefill processes the prompt; decode generates tokens while reusing cached keys and values.
  • Paged KV storage, prefix caching and continuous batching improve capacity and utilization.
  • Quantization and speculative decoding need workload-specific speed and quality validation.
  • Measure queueing, first-token and per-token latency under realistic concurrency.
  • Tune routing, parallelism and deployment around goodput and explicit service targets.

Key Acronyms, Concepts & Jargon

TermShort meaning
KV cacheKey-Value cache: saves attention state so earlier tokens need not be recomputed.
Prefill / decodePrompt processing / subsequent token generation.
PagedAttentionStores KV state in blocks to reduce wasted memory and support sharing.
Continuous batchingAdds and removes requests as generation proceeds.
Prefix cachingReuses cached computation for an identical leading token sequence.
QuantizationStores or computes values at lower precision to reduce resource use.
GPTQ / AWQPost-training weight quantization / Activation-aware Weight Quantization methods.
Speculative decodingProposes tokens cheaply, then verifies them with the target model.
TTFT / TPOT / ITLTime To First Token / Time Per Output Token / Inter-Token Latency.
Throughput / goodputWork completed per time / work completed within the specified targets.
Concurrency / batch sizeRequests in flight / requests processed together in a model step.
p95 / p99Latency thresholds below which 95% / 99% of observations fall.
SLO / SSEService Level Objective / Server-Sent Events: service target / HTTP streaming format.
Disaggregated servingRuns prefill and decode on separate resources.

Back to section overview

⚡AI-assisted content - always verify, always explore multiple perspectives·