Contents
Map

08 · Inference & Serving

FastAPI + vLLM Endpoint

View as:

Code Lab 01 - FastAPI + vLLM Endpoint

A real FastAPI endpoint backed by a vLLM V1 AsyncLLM, serving a small open instruction model with both streaming (SSE) and non-streaming responses - plus a concurrent load-testing client that measures tokens/sec, p50/p95 latency, and time-to-first-token under load.

← Back to Overview: Serving & Inference · Back to Concepts: vLLM & Paged Attention · Batching, Concurrency & Latency · Streaming & TTFT

Learning objectives 2 hours (plus GPU time)
By the end of this page you will be able to:
  • Serve a model with vLLM's AsyncLLM behind a FastAPI endpoint with streaming (SSE) and non-streaming responses
  • Load-test the endpoint and report tokens per second, p50/p95 latency and TTFT at several concurrency levels
  • Find the concurrency at which a latency SLO breaks, and explain it from batching and KV-cache behaviour
Prerequisites

What's In This Lab

PropertyDetail
TaskServe a small open instruction model behind a production-shaped HTTP endpoint
Base modelA small open instruction model (default: Qwen/Qwen3-1.7B) - the same class of model used in Fine-Tuning Lab; a QLoRA-merged model from that lab works equally well as a drop-in --model swap
Serving enginevLLM V1 AsyncLLM (Paged Attention + continuous batching)
Verifiedserve.py compiles and its vLLM imports and calls (AsyncEngineArgs, AsyncLLM.from_engine_args, generate, shutdown) were checked against vLLM's source and API docs in September 2026 (vLLM ≥ 0.11, V1 engine only). It has not been run on a GPU as part of this review.
Requests{"prompt": "..."} for raw text, or {"messages": [...]} to apply the model's chat template (use this for instruct models)
API layerFastAPI, /generate endpoint supporting stream=true (SSE) and stream=false (single JSON response)
Load testConcurrent async client measuring tokens/sec, p50/p95 latency, and TTFT at configurable concurrency
ComplexityIntermediate-Advanced
Files01-FastAPI-vLLM-Endpoint/{serve.py, load_test.py, requirements.txt, README.mdx}

Architecture

flowchart TD
    subgraph Serve["serve.py"]
        direction LR
        Engine["AsyncLLM (V1)\n(vLLM, Paged Attention +\ncontinuous batching)"] --> API["FastAPI /generate\nstream=true/false"]
        API --> SSE["StreamingResponse\n(SSE token-by-token)"]
        API --> JSON["Single JSON response\n(non-streaming)"]
    end
    subgraph LoadTest["load_test.py"]
        direction LR
        Fire["Fire N concurrent\nasync requests"] --> Measure["Measure TTFT +\nper-request latency"]
        Measure --> Report["Report tok/s,\np50/p95 latency, TTFT"]
    end

    SSE --> Fire
    JSON --> Fire

    style Engine fill:#e8e0d4,stroke:#c8b89a
    style API fill:#dde4dc,stroke:#b0c4b0
    style SSE fill:#d8dfe8,stroke:#b0bac8
    style JSON fill:#d8dfe8,stroke:#b0bac8
    style Fire fill:#ddd8e4,stroke:#b8b0c8
    style Measure fill:#ddd8e4,stroke:#b8b0c8
    style Report fill:#e8e0d4,stroke:#c8b89a

Running

cd 08-Serving-and-Inference/CodeLabs/01-FastAPI-vLLM-Endpoint
pip install -r requirements.txt

# Terminal 1: start the server
python serve.py --model Qwen/Qwen3-1.7B --port 8000

# Terminal 2: run the load test against it
python load_test.py --url http://localhost:8000/generate --concurrency 16 --requests 64

Full details, code walkthrough, and what each part demonstrates: see README.

Check Yourself

Check yourself
0 / 3 answered
  1. Why should the load test report p95 latency rather than the average?
  2. TTFT rises sharply at concurrency 64 while aggregate tokens/s still grows. What is happening?
  3. Why does the server use AsyncLLM rather than calling LLM.generate() per request?

Exercises

Exercise - Map the latency curve

Run load_test.py at concurrency 1, 4, 16, 64 and 128 with a fixed prompt set. Plot p50/p95 TTFT, p95 end-to-end latency and aggregate tokens/s. For an SLO of p95 TTFT < 500 ms, what capacity do you report?

Solution

Report the highest concurrency where p95 TTFT stays under 500 ms, minus a safety margin; show that aggregate throughput keeps rising past that point, which is why throughput alone is the wrong capacity metric.

Exercise - Prefix caching on and off

Give every request the same 1,500-token system prompt. Run the load test with prefix caching enabled and disabled (enable_prefix_caching in the engine args) and compare TTFT and throughput.

Solution

With caching, the shared prefix is prefilled once and reused, so TTFT drops sharply and throughput rises because prefill compute is freed for decoding; without it, each request pays the full prefill.

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·