GPT-5.6 Luna vs Gemini 2.5 Flash-Lite
Gemini 2.5 Flash-Lite is the cheaper of the two across every workload shape tested below. Headline rates are $0.2 in / $1.2 out for GPT-5.6 Luna versus $0.1 in / $0.4 out for Gemini 2.5 Flash-Lite, per million tokens.
Real monthly cost, three workloads
| Workload | GPT-5.6 Luna | Gemini 2.5 Flash-Lite | Cheaper |
|---|---|---|---|
| Support chatbot 1.5k in / 400 out · 100k req/mo · 30% cached | $69.90 | $26.95 | Gemini 2.5 Flash-Lite (2.6× cheaper) |
| RAG / doc Q&A 12k in / 700 out · 50k req/mo · 60% cached | $97.20 | $41.60 | Gemini 2.5 Flash-Lite (2.3× cheaper) |
| Coding agent 25k in / 2.5k out · 20k req/mo · 80% cached | $88.00 | $34.00 | Gemini 2.5 Flash-Lite (2.6× cheaper) |
Costs include prompt caching at the stated hit rate and per-model tokenizer correction. Adjust for your own traffic →
Specifications side by side
| GPT-5.6 Luna | Gemini 2.5 Flash-Lite | |
|---|---|---|
| Provider | OpenAI | |
| Input / M tokens | $0.2 | $0.1 |
| Output / M tokens | $1.2 | $0.4 |
| Cached input | $0.02 | $0.01 |
| Context window | 1050K | — |
| Batch discount | −50% | −50% |
| Tokenizer factor | ×1 | ×1 |
| Min prefix to cache | — | — |
| Announced retirement | — | — |
Which should you pick?
On cost alone, Gemini 2.5 Flash-Lite wins every workload shape above.
Full detail: GPT-5.6 Luna · Gemini 2.5 Flash-Lite
All figures verified 2026-08-18 against official provider documentation. Fields a provider does not publish are shown as “—” rather than estimated.