Prompt Cache Analyzer

Paste your system prompt to find out whether it is actually long enough to cache. Anthropic silently ignores cache_control when the prefix falls below a per-model minimum — no error, no warning. You pay full price on every request while believing caching is on.

Runs in your browser · your prompt is never uploaded · thresholds verified 2026-08-18

Stable prefix — system prompt, tool schemas, fixed context
The part that is byte-identical on every request. This is what can be cached.
Variable part — the user turn
Changes every request, so it is always billed at full price.
Requests / month
Cache TTL
Prefix reuse rate — 85%
Prefix tokens
…
ModelMin to cacheYour prefixCaches?Monthly, no cacheMonthly, cachedSaved

Loading tokenizer…

Minimum cacheable prefix, by model

The spread is eight-fold. This is the single most overlooked number in LLM cost work, because falling below it produces no signal of any kind.

Model Minimum tokens to cache
claude-opus-5 512
claude-fable-5 512
claude-mythos-5 512
claude-opus-4-8 1,024
claude-sonnet-5 1,024
claude-sonnet-4-6 1,024
claude-sonnet-4-5 1,024
claude-opus-4-7 2,048
claude-opus-4-6 4,096
claude-opus-4-5 4,096
claude-haiku-4-5 4,096

Only Anthropic publishes explicit thresholds, so only Anthropic models are analysed here. We do not guess at other providers' behaviour.

Frequently asked questions

What is the minimum prompt length for caching to work?

It varies by model and the range is wide. Claude Opus 5, Fable 5 and Mythos 5 cache from 512 tokens. Claude Sonnet 5, Sonnet 4.6 and Sonnet 4.5 need 1,024. Claude Opus 4.8 needs 1,024, Opus 4.7 needs 2,048, and Claude Opus 4.6, Opus 4.5 and Haiku 4.5 need 4,096 — eight times the Opus 5 threshold. A prompt that caches efficiently on one model can silently fail to cache on another.

What happens if my prompt is below the minimum?

Nothing visible. The request succeeds, no error is raised, and cache_control is simply ignored. You pay the full input price on every single request while believing caching is active. The only way to detect it is to check that both cache_creation_input_tokens and cache_read_input_tokens are greater than zero in the response usage — which most teams never look at.

How many times must a prefix be reused before caching pays off?

A 5-minute cache costs 1.25× to write and 0.1× to read, so it breaks even after a single reuse. A 1-hour cache costs 2× to write, so it needs two reuses. Below that, caching is actively more expensive than not caching — which matters for low-traffic endpoints where the cache expires between requests.

What invalidates a prompt cache?

The cache follows a hierarchy: tools → system → messages. Changing tool definitions invalidates everything downstream. Toggling web search or citations modifies the system prompt and invalidates system and message caches. Switching between fast and standard speed does the same. Adding or removing images affects message blocks. The practical rule: place cache_control on the last block whose prefix is byte-identical across requests, and never on anything containing a timestamp.

How many cache breakpoints can I use?

Up to four explicit breakpoints per request, which lets you cache sections that change at different frequencies — a never-changing tool schema, a daily-rotating knowledge base, a per-conversation history. Automatic caching consumes one of the four slots.

Is caching or switching to a cheaper model the bigger saving?

For any workload that re-sends a substantial fixed prefix — agents, RAG, classification with few-shot examples — caching is usually the larger lever. A cache read costs 10% of base input price, a 90% reduction on the cached portion. Moving from a mid-tier to a budget model rarely achieves that, and costs you capability. Try caching first.

Source

Anthropic — prompt caching documentation