DeepSeek V4 Flash pricing
DeepSeek V4 Flash costs $0.44 per million input tokens and $1.32 per million output tokens , with cached input at $0.014. Its context window is 1M tokens. Rates shown are peak; off-peak hours cost $0.22 in / $0.66 out.
Price per million tokens
| Input | $0.44 |
| Cached input | $0.014 |
| Output | $1.32 |
| Context window | 1M tokens |
| Batch discount | not offered |
Time-of-day pricing: 01:00-04:00 and 06:00-10:00 UTC are peak; all other hours bill at half
What it costs in practice
Monthly cost for three common workload shapes, with prompt caching applied.
| Workload | Shape | Per request | Per month |
|---|---|---|---|
| Support chatbot | 1.5k in / 400 out, 100k requests/month, 30% cached | $0.000996 | $99.63 |
| RAG / document Q&A | 12k in / 700 out, 50k requests/month, 60% cached | $0.00314 | $156.84 |
| Coding agent | 25k in / 2.5k out, 20k requests/month, 80% cached | $0.00578 | $115.60 |
Model your own workload with caching, batching and tokenizer correction →
Comparable models from other providers
| Model | Provider | Input | Output |
|---|---|---|---|
| GPT-4.1 mini | OpenAI | $0.4 | $1.6 |
| Gemini 3.5 Flash-Lite | $0.3 | $2.5 | |
| Gemini 2.5 Flash | $0.3 | $2.5 | |
| GPT-5.4 mini | OpenAI | $0.75 | $4.5 |
All figures verified 2026-08-18 against DeepSeek's official pricing documentation . Fields the provider does not publish are shown as “—” rather than estimated.