Prompt Cache Analyzer
Paste your system prompt to find out whether it is actually long enough to cache.
Anthropic silently ignores cache_control when the
prefix falls below a per-model minimum — no error, no warning. You pay full price on
every request while believing caching is on.
Runs in your browser · your prompt is never uploaded · thresholds verified 2026-08-18
| Model | Min to cache | Your prefix | Caches? | Monthly, no cache | Monthly, cached | Saved |
|---|
Loading tokenizer…
Minimum cacheable prefix, by model
The spread is eight-fold. This is the single most overlooked number in LLM cost work, because falling below it produces no signal of any kind.
| Model | Minimum tokens to cache |
|---|---|
| claude-opus-5 | 512 |
| claude-fable-5 | 512 |
| claude-mythos-5 | 512 |
| claude-opus-4-8 | 1,024 |
| claude-sonnet-5 | 1,024 |
| claude-sonnet-4-6 | 1,024 |
| claude-sonnet-4-5 | 1,024 |
| claude-opus-4-7 | 2,048 |
| claude-opus-4-6 | 4,096 |
| claude-opus-4-5 | 4,096 |
| claude-haiku-4-5 | 4,096 |
Only Anthropic publishes explicit thresholds, so only Anthropic models are analysed here. We do not guess at other providers' behaviour.
Frequently asked questions
What is the minimum prompt length for caching to work?
It varies by model and the range is wide. Claude Opus 5, Fable 5 and Mythos 5 cache from 512 tokens. Claude Sonnet 5, Sonnet 4.6 and Sonnet 4.5 need 1,024. Claude Opus 4.8 needs 1,024, Opus 4.7 needs 2,048, and Claude Opus 4.6, Opus 4.5 and Haiku 4.5 need 4,096 — eight times the Opus 5 threshold. A prompt that caches efficiently on one model can silently fail to cache on another.
What happens if my prompt is below the minimum?
Nothing visible. The request succeeds, no error is raised, and cache_control is simply ignored. You pay the full input price on every single request while believing caching is active. The only way to detect it is to check that both cache_creation_input_tokens and cache_read_input_tokens are greater than zero in the response usage — which most teams never look at.
How many times must a prefix be reused before caching pays off?
A 5-minute cache costs 1.25× to write and 0.1× to read, so it breaks even after a single reuse. A 1-hour cache costs 2× to write, so it needs two reuses. Below that, caching is actively more expensive than not caching — which matters for low-traffic endpoints where the cache expires between requests.
What invalidates a prompt cache?
The cache follows a hierarchy: tools → system → messages. Changing tool definitions invalidates everything downstream. Toggling web search or citations modifies the system prompt and invalidates system and message caches. Switching between fast and standard speed does the same. Adding or removing images affects message blocks. The practical rule: place cache_control on the last block whose prefix is byte-identical across requests, and never on anything containing a timestamp.
How many cache breakpoints can I use?
Up to four explicit breakpoints per request, which lets you cache sections that change at different frequencies — a never-changing tool schema, a daily-rotating knowledge base, a per-conversation history. Automatic caching consumes one of the four slots.
Is caching or switching to a cheaper model the bigger saving?
For any workload that re-sends a substantial fixed prefix — agents, RAG, classification with few-shot examples — caching is usually the larger lever. A cache read costs 10% of base input price, a 90% reduction on the cached portion. Moving from a mid-tier to a budget model rarely achieves that, and costs you capability. Try caching first.
Source
Anthropic — prompt caching documentation