GPT-5 mini vs DeepSeek V4 Flash
DeepSeek V4 Flash is the cheaper of the two across most of the workload shapes tested below. Headline rates are $0.25 in / $2 out for GPT-5 mini versus $0.44 in / $1.32 out for DeepSeek V4 Flash, per million tokens.
Real monthly cost, three workloads
| Workload | GPT-5 mini | DeepSeek V4 Flash | Cheaper |
|---|---|---|---|
| Support chatbot 1.5k in / 400 out · 100k req/mo · 30% cached | $107.38 | $99.63 | DeepSeek V4 Flash (1.1× cheaper) |
| RAG / doc Q&A 12k in / 700 out · 50k req/mo · 60% cached | $139.00 | $156.84 | GPT-5 mini (1.1× cheaper) |
| Coding agent 25k in / 2.5k out · 20k req/mo · 80% cached | $135.00 | $115.60 | DeepSeek V4 Flash (1.2× cheaper) |
Costs include prompt caching at the stated hit rate and per-model tokenizer correction. Adjust for your own traffic →
Specifications side by side
| GPT-5 mini | DeepSeek V4 Flash | |
|---|---|---|
| Provider | OpenAI | DeepSeek |
| Input / M tokens | $0.25 | $0.44 |
| Output / M tokens | $2 | $1.32 |
| Cached input | $0.025 | $0.014 |
| Context window | — | 1000K |
| Batch discount | −50% | — |
| Tokenizer factor | ×1 | ×1 |
| Min prefix to cache | — | — |
| Announced retirement | — | — |
Which should you pick?
On cost alone, DeepSeek V4 Flash wins the majority of the shapes above. The ranking flips depending on the input-to-output ratio, so check the shape closest to your own traffic.
Full detail: GPT-5 mini · DeepSeek V4 Flash
All figures verified 2026-08-18 against official provider documentation. Fields a provider does not publish are shown as “—” rather than estimated.