Claude Sonnet 5 vs DeepSeek V4 Pro
DeepSeek V4 Pro is the cheaper of the two across every workload shape tested below. Headline rates are $2 in / $10 out for Claude Sonnet 5 versus $1.32 in / $3.96 out for DeepSeek V4 Pro, per million tokens. Note that sticker price understates the gap: Claude Sonnet 5 tokenizes the same text into roughly 30% more tokens, which the figures below account for.
Real monthly cost, three workloads
| Workload | Claude Sonnet 5 | DeepSeek V4 Pro | Cheaper |
|---|---|---|---|
| Support chatbot 1.5k in / 400 out · 100k req/mo · 30% cached | $804.70 | $298.98 | DeepSeek V4 Pro (2.7× cheaper) |
| RAG / doc Q&A 12k in / 700 out · 50k req/mo · 60% cached | $1,173 | $471.24 | DeepSeek V4 Pro (2.5× cheaper) |
| Coding agent 25k in / 2.5k out · 20k req/mo · 80% cached | $1,014 | $347.60 | DeepSeek V4 Pro (2.9× cheaper) |
Costs include prompt caching at the stated hit rate and per-model tokenizer correction. Adjust for your own traffic →
Specifications side by side
| Claude Sonnet 5 | DeepSeek V4 Pro | |
|---|---|---|
| Provider | Anthropic | DeepSeek |
| Input / M tokens | $2 | $1.32 |
| Output / M tokens | $10 | $3.96 |
| Cached input | $0.2 | $0.044 |
| Context window | 1000K | 1000K |
| Batch discount | −50% | — |
| Tokenizer factor | ×1.3 | ×1 |
| Min prefix to cache | 1024 | — |
| Announced retirement | 2027-06-30+ | — |
Which should you pick?
On cost alone, DeepSeek V4 Pro wins every workload shape above.
Caching behaviour differs. Claude Sonnet 5 requires a prefix of at least 1,024 tokens before caching activates.
Below those thresholds cache_control is ignored silently, with no error — which can quietly
erase the saving you were counting on.
Lifecycle matters as much as price here. Claude Sonnet 5 has an announced shutdown date of 2027-06-30 or later. A model that is marginally cheaper but retires within your planning horizon is rarely the better choice.
Full detail: Claude Sonnet 5 · DeepSeek V4 Pro
All figures verified 2026-08-18 against official provider documentation. Fields a provider does not publish are shown as “—” rather than estimated.