Code Lab 01 - FastAPI + vLLM Endpoint
A real FastAPI endpoint backed by a vLLM V1 AsyncLLM, serving a small open instruction model with both streaming (SSE) and non-streaming responses - plus a concurrent load-testing client that measures tokens/sec, p50/p95 latency, and time-to-first-token under load.
← Back to Overview: Serving & Inference · Back to Concepts: vLLM & Paged Attention · Batching, Concurrency & Latency · Streaming & TTFT
- Serve a model with vLLM's AsyncLLM behind a FastAPI endpoint with streaming (SSE) and non-streaming responses
- Load-test the endpoint and report tokens per second, p50/p95 latency and TTFT at several concurrency levels
- Find the concurrency at which a latency SLO breaks, and explain it from batching and KV-cache behaviour
- vLLM & Paged Attention, Batching, Concurrency & Latency and Streaming & TTFT
- A CUDA GPU with at least 16 GB for the default 1.7B model
What's In This Lab
| Property | Detail |
|---|---|
| Task | Serve a small open instruction model behind a production-shaped HTTP endpoint |
| Base model | A small open instruction model (default: Qwen/Qwen3-1.7B) - the same class of model used in Fine-Tuning Lab; a QLoRA-merged model from that lab works equally well as a drop-in --model swap |
| Serving engine | vLLM V1 AsyncLLM (Paged Attention + continuous batching) |
| Verified | serve.py compiles and its vLLM imports and calls (AsyncEngineArgs, AsyncLLM.from_engine_args, generate, shutdown) were checked against vLLM's source and API docs in September 2026 (vLLM ≥ 0.11, V1 engine only). It has not been run on a GPU as part of this review. |
| Requests | {"prompt": "..."} for raw text, or {"messages": [...]} to apply the model's chat template (use this for instruct models) |
| API layer | FastAPI, /generate endpoint supporting stream=true (SSE) and stream=false (single JSON response) |
| Load test | Concurrent async client measuring tokens/sec, p50/p95 latency, and TTFT at configurable concurrency |
| Complexity | Intermediate-Advanced |
| Files | 01-FastAPI-vLLM-Endpoint/{serve.py, load_test.py, requirements.txt, README.mdx} |
Architecture
flowchart TD
subgraph Serve["serve.py"]
direction LR
Engine["AsyncLLM (V1)\n(vLLM, Paged Attention +\ncontinuous batching)"] --> API["FastAPI /generate\nstream=true/false"]
API --> SSE["StreamingResponse\n(SSE token-by-token)"]
API --> JSON["Single JSON response\n(non-streaming)"]
end
subgraph LoadTest["load_test.py"]
direction LR
Fire["Fire N concurrent\nasync requests"] --> Measure["Measure TTFT +\nper-request latency"]
Measure --> Report["Report tok/s,\np50/p95 latency, TTFT"]
end
SSE --> Fire
JSON --> Fire
style Engine fill:#e8e0d4,stroke:#c8b89a
style API fill:#dde4dc,stroke:#b0c4b0
style SSE fill:#d8dfe8,stroke:#b0bac8
style JSON fill:#d8dfe8,stroke:#b0bac8
style Fire fill:#ddd8e4,stroke:#b8b0c8
style Measure fill:#ddd8e4,stroke:#b8b0c8
style Report fill:#e8e0d4,stroke:#c8b89a
Running
cd 08-Serving-and-Inference/CodeLabs/01-FastAPI-vLLM-Endpoint
pip install -r requirements.txt
# Terminal 1: start the server
python serve.py --model Qwen/Qwen3-1.7B --port 8000
# Terminal 2: run the load test against it
python load_test.py --url http://localhost:8000/generate --concurrency 16 --requests 64
Full details, code walkthrough, and what each part demonstrates: see README.
Check Yourself
- Why should the load test report p95 latency rather than the average?
- TTFT rises sharply at concurrency 64 while aggregate tokens/s still grows. What is happening?
- Why does the server use AsyncLLM rather than calling LLM.generate() per request?
Exercises
Run load_test.py at concurrency 1, 4, 16, 64 and 128 with a fixed prompt set. Plot p50/p95 TTFT, p95 end-to-end latency and aggregate tokens/s. For an SLO of p95 TTFT < 500 ms, what capacity do you report?
Solution
Report the highest concurrency where p95 TTFT stays under 500 ms, minus a safety margin; show that aggregate throughput keeps rising past that point, which is why throughput alone is the wrong capacity metric.
Give every request the same 1,500-token system prompt. Run the load test with prefix caching enabled and disabled (enable_prefix_caching in the engine args) and compare TTFT and throughput.
Solution
With caching, the shared prefix is prefilled once and reused, so TTFT drops sharply and throughput rises because prefill compute is freed for decoding; without it, each request pays the full prefill.
References
- Kwon et al., PagedAttention (vLLM) (2023)
- vLLM project, vLLM documentation - engine arguments, OpenAI-compatible server, benchmarking (2026)
- FastAPI, StreamingResponse (2026)
Last reviewed: 2026-09