Claude Sonnet 5 pricing

Claude Sonnet 5 costs $2 per million input tokens and $10 per million output tokens , with cached input at $0.2. Its context window is 1M tokens. Batch requests are discounted 50%.

Retirement scheduled. Anthropic lists this model for shutdown on 2027-06-30 (a "not sooner than" floor — at least 60 days of notice precedes the real date). See all shutdown dates →

Price per million tokens

Input $2
Cached input $0.2
Output $10
Context window 1M tokens
Batch discount −50%
Minimum prefix to cache 1,024 tokens
Tokenizer factor ×1.3

The $2/$10 introductory rate is now standard; the planned 1 Sep rise was cancelled

What it costs in practice

Monthly cost for three common workload shapes, with prompt caching applied.

Workload Shape Per request Per month
Support chatbot 1.5k in / 400 out, 100k requests/month, 30% cached $0.00805 $804.70
RAG / document Q&A 12k in / 700 out, 50k requests/month, 60% cached $0.0235 $1,173
Coding agent 25k in / 2.5k out, 20k requests/month, 80% cached $0.0507 $1,014

Model your own workload with caching, batching and tokenizer correction →

Caching on Claude Sonnet 5

A cached prefix on this model must be at least 1,024 tokens. Below that, Anthropic ignores cache_control silently — no error is returned and you pay full input price on every request. A cache hit costs $0.2 per million tokens, roughly 10% of the base input rate.

Check whether your system prompt clears the threshold →

Comparable models from other providers

Model Provider Input Output
GPT-5.6 Terra OpenAI $2 $12
GPT-4.1 OpenAI $2 $8
o3 OpenAI $2 $8
Gemini 3.1 Pro (preview) Google $2 $12

All figures verified 2026-08-18 against Anthropic's official pricing documentation . Fields the provider does not publish are shown as “—” rather than estimated.