Prompt Caching & Cost
Prompt caching reuses the model's already-computed state for a repeated prompt prefix, cutting the price of those input tokens by up to about 90% and time-to-first-token substantially - if the prompt is structured so the prefix actually repeats.
- Explain what a prompt cache stores and why only an identical prefix can hit it
- Compare the caching models of the major APIs and self-hosted engines (automatic vs explicit, TTLs, minimums, prices)
- Structure prompts and agent loops for high cache hit rates, and compute the break-even for an explicit cache write
- Combine caching with the other cost levers - batch APIs, model routing, effort and output length
- KV Cache & Inference Optimization - what the KV cache is
- Context Engineering
What Gets Cached
Processing the input (prefill) computes key and value tensors for every token at every layer - the KV cache. Because attention is causal, the KV entries for a token depend only on the tokens before it. So if a new request starts with exactly the same tokens as an earlier one, the server can reuse the stored KV entries for that prefix and only compute the new suffix.
flowchart LR
subgraph R1["Request 1"]
A1["โ๏ธ System prompt"] --> B1["๐ง Tools"] --> C1["๐ Document"] --> D1["โ Question A"]
end
subgraph R2["Request 2"]
A2["โ๏ธ System prompt"] --> B2["๐ง Tools"] --> C2["๐ Document"] --> D2["โ Question B"]
end
C1 -.->|"identical prefix:<br/>KV reused (cache hit)"| C2
D2 --> N(["only Question B<br/>is computed"])
style A2 fill:#dde4dc,stroke:#b0c4b0
style B2 fill:#dde4dc,stroke:#b0c4b0
style C2 fill:#dde4dc,stroke:#b0c4b0
style D2 fill:#e8e0d4,stroke:#c8b89a
Consequences:
- Only prefixes match. One changed token invalidates everything after it. A timestamp at the top of the system prompt defeats the cache for the whole request.
- Order matters. Put stable content first (instructions, tool definitions, long documents, examples), then conversation history, then the new request.
- Serialisation must be byte-stable. Tool definitions in a different order, JSON with unsorted keys, or whitespace differences all change the tokens.
- Outputs are not cached - this is not a response cache. The model still generates fresh output every time.
Self-hosted engines do the same thing automatically: vLLM's automatic prefix caching hashes KV-cache blocks, and SGLang's RadixAttention keeps prefixes in a radix tree (see vLLM & PagedAttention).
How the Major APIs Do It
| Anthropic | OpenAI | Google Gemini | |
|---|---|---|---|
| Activation | Explicit cache_control breakpoints on content blocks (up to 4), plus an automatic mode that moves a breakpoint along the growing conversation | Automatic for supported models; optional prompt_cache_key to group requests | Implicit caching automatic on Gemini 2.5 and later; explicit context caches you create with a TTL |
| Minimum prefix | Model-dependent (512 to 4,096 tokens) | 1,024 tokens | 2,048 (2.5 models) or 4,096 (3.x models) for implicit caching |
| Lifetime | 5 minutes (default) or 1 hour, refreshed on each hit | In-memory, typically minutes; newer models guarantee at least 30 minutes after last use; some models offer 24-hour retention | Implicit: short, best effort; explicit: the TTL you set |
| Price of a hit | ~0.1ร the input price (lower on some models) | Up to 90% off (0.1ร on newest models) | Discounted - see the pricing page per model |
| Price of a write | 1.25ร input (5-minute) or 2ร (1-hour) | No surcharge | Implicit: none; explicit: storage charged per hour |
| Where to see it | usage.cache_read_input_tokens, cache_creation_input_tokens | usage.prompt_tokens_details.cached_tokens (Responses: input_tokens_details) | Cached-token count in the usage metadata |
Details change often - check each provider's caching page before relying on a number. The design rules below don't change.
Break-even for explicit writes
With Anthropic's 5-minute cache, a write costs 1.25ร and each read 0.1ร the base input price. For a prefix reused n times within the TTL:
uncached cost = n ร 1.0
cached cost = 1.25 + (n โ 1) ร 0.1
break-even: n = 2 โ 2.0 vs 1.35 (caching already wins at the second request)
1-hour TTL (2ร write): n = 3 โ 3.0 vs 2.2
A read refreshes the TTL, so steady traffic keeps a 5-minute cache warm indefinitely; the 1-hour TTL is for gaps between 5 and 60 minutes.
Designing for Cache Hits
import anthropic
client = anthropic.Anthropic()
SYSTEM = [
{"type": "text", "text": LONG_STABLE_INSTRUCTIONS},
{"type": "text", "text": PRODUCT_HANDBOOK, # large, shared by every request
"cache_control": {"type": "ephemeral"}}, # breakpoint: cache everything up to here
]
def answer(question: str, history: list[dict]):
return client.messages.create(
model=MODEL, max_tokens=2048,
system=SYSTEM, # identical bytes on every call
messages=history + [{"role": "user", "content": question}],
)
# resp.usage.cache_read_input_tokens shows the tokens served from cache
Checklist:
- Stable first, variable last: instructions โ tools โ documents โ examples โ history โ current input.
- No per-request values in the prefix: move dates, user names and request IDs into the final user message.
- Freeze tool definitions and their order; loading tools dynamically at the top of the prompt breaks the cache on every change.
- Append, don't rewrite, history in agent loops - editing or summarising early turns invalidates everything after the edit. Compact rarely and deliberately (see Context Engineering).
- Measure the hit rate from the usage fields; a sudden drop usually means something changed in the prefix.
The Other Cost Levers
| Lever | Typical effect | Trade-off |
|---|---|---|
| Prompt caching | Up to ~90% off repeated input; lower TTFT | Prompt structure constraints |
| Batch APIs (OpenAI, Anthropic and Google all offer one) | About 50% off, results within 24 hours | Not for interactive traffic |
| Model routing / cascades | Send easy requests to a small model, escalate hard ones | Needs a router and an eval per route |
| Reasoning effort | Lower effort cuts output tokens sharply on easy tasks | Accuracy on hard tasks - measure per task |
| Output length | Output tokens cost several times more than input tokens | Ask for concise formats; enforce with schemas and max_tokens |
| Context size | Fewer retrieved chunks, trimmed tool results | Recall - see RAG evaluation |
A simple cost model per request:
cost = uncached_input ร p_in + cached_input ร p_cache + output (incl. reasoning) ร p_out
Log the three token counts for every call. Most cost surprises come from output (reasoning) tokens and from cache hit rates that quietly dropped after a prompt change.
Check Yourself
- You add the current date and time to the first line of your system prompt. What happens to prompt caching?
- With a 1.25ร write price and 0.1ร read price, how many requests sharing a prefix within the TTL are needed for caching to be cheaper than no caching?
- What does a prompt cache store?
- An agent's cache hit rate fell from 85% to 10% after a release. Name two likely causes to check first.
Exercises
An assistant sends a 6,000-token system prompt + handbook, 1,500 tokens of retrieved chunks, 500 tokens of history and a 100-token question, and returns 400 output tokens. Prices: input $3/M, cached read $0.30/M, cache write 1.25ร, output $15/M. Traffic is continuous. Compute the cost per request with and without caching the 6,000-token prefix.
Solution
Without caching: 8,100 ร $3/M + 400 ร $15/M = $0.0243 + $0.0060 = $0.0303. With caching (steady state, all hits): 6,000 ร $0.30/M + 2,100 ร $3/M + $0.0060 = $0.0018 + $0.0063 + $0.0060 = $0.0141 - about 53% cheaper. Output is now over 40% of the bill, so output length is the next lever.
This prompt order is used for every call: [current date] [user profile] [system rules] [tool definitions] [handbook] [history] [question]. Reorder it for caching and say where the date and profile should go.
Solution
[system rules] [tool definitions] [handbook] (cache breakpoint) [history] [user profile + current date + question]. The profile and date vary per user or per call, so they belong after the cached prefix - typically in the final user message.
Study Notes
Must-know:
- Caching reuses prefix KV state; only exact token prefixes hit; outputs are never cached
- Anthropic: explicit breakpoints, 5 min / 1 h TTL, write 1.25ร / 2ร, read ~0.1ร; OpenAI: automatic, โฅ1,024 tokens, up to 90% off; Gemini: implicit on 2.5+, plus explicit caches
- Stable content first; no timestamps or IDs in the prefix; frozen tool order; append-only history
- Break-even for a 5-minute write is two requests
- Other levers: batch (~50%), routing, effort, output length, smaller contexts - track uncached, cached and output tokens separately
References
- Anthropic, Prompt caching
- OpenAI, Prompt caching guide
- Google, Gemini context caching
- Gim et al., Prompt Cache: Modular Attention Reuse for Low-Latency Inference (MLSys 2024)
- Zheng et al., SGLang: Efficient Execution of Structured Language Model Programs (NeurIPS 2024) - RadixAttention
- vLLM, Automatic Prefix Caching
Last reviewed: 2026-09