Long-Context Price Cliff Calculator
6 of the models tracked here switch to a higher rate once a prompt passes 200,000 tokens. The higher rate is not charged on the overflow — it is charged on every token in the request, output included. On Grok 4.6, a prompt just under 200,000 tokens and one just over it differ in price by 2.00×. One token.
Runs in your browser · tier pricing verified 2026-08-18 against official provider docs
| Model | Provider | Context window | Cliff at | $/M in (applied) | Monthly here | Monthly past the cliff |
|---|---|---|---|---|---|---|
| GPT-5 nano | OpenAI | — | none | 0.05 | $21.10 | flat |
| Gemini 2.5 Flash-Lite | — | none | 0.1 | $40.60 | flat | |
| GPT-4o mini | OpenAI | — | none | 0.15 | $60.90 | flat |
| GPT-5.6 Luna | OpenAI | 1.1M | none | 0.2 | $82.80 | flat |
| GPT-5.4 nano | OpenAI | — | none | 0.2 | $83.00 | flat |
| Gemini 3.1 Flash-Lite | — | none | 0.25 | $103.50 | flat | |
| GPT-5 mini | OpenAI | — | none | 0.25 | $105.50 | flat |
| Gemini 3.5 Flash-Lite | — | none | 0.3 | $127.00 | flat | |
| Gemini 2.5 Flash | 1M | none | 0.3 | $127.00 | flat | |
| GPT-4.1 mini | OpenAI | — | none | 0.4 | $162.40 | flat |
| DeepSeek V4 Flash | DeepSeek | 1M | none | 0.44 | $176.88 | flat |
| Gemini 3.7 Flash | — | none | 0.75 | $307.50 | flat | |
| Gemini 3.6 Flash | — | none | 0.75 | $307.50 | flat | |
| GPT-5.4 mini | OpenAI | — | none | 0.75 | $310.50 | flat |
| Grok Build 0.1 | xAI | 256K | 200K | 1 | $398.00 | $816.00 ×2.05 |
| Claude Haiku 4.5 | Anthropic | 200K | none | 1 | $410.00 | flat |
| o4-mini | OpenAI | — | none | 1.1 | $446.60 | flat |
| Grok 4.3 | xAI | 1M | 200K | 1.25 | $497.50 | $1,020 ×2.05 |
| GPT-5.1 | OpenAI | — | none | 1.25 | $527.50 | flat |
| GPT-5 | OpenAI | — | none | 1.25 | $527.50 | flat |
| Gemini 2.5 Pro | — | 200K | 1.25 | $527.50 | $1,060 ×2.01 | |
| DeepSeek V4 Pro | DeepSeek | 1M | none | 1.32 | $530.64 | flat |
| Gemini 3.5 Flash | — | none | 1.5 | $621.00 | flat | |
| GPT-5.3 Codex | OpenAI | — | none | 1.75 | $738.50 | flat |
| GPT-5.2 | OpenAI | — | none | 1.75 | $738.50 | flat |
| Grok 4.6 | xAI | 500K | 200K | 2 | $804.00 | $1,648 ×2.05 |
| Grok 4.5 | xAI | 500K | 200K | 2 | $804.00 | $1,648 ×2.05 |
| GPT-4.1 | OpenAI | — | none | 2 | $812.00 | flat |
| o3 | OpenAI | 200K | none | 2 | $812.00 | flat |
| GPT-5.6 Terra | OpenAI | 1.1M | none | 2 | $828.00 | flat |
| Gemini 3.1 Pro (preview) | — | 200K | 2 | $828.00 | $1,672 ×2.02 | |
| GPT-5.4 | OpenAI | 272K | none | 2.5 | $1,035 | flat |
| Claude Sonnet 5 | Anthropic | 1M | none | 2 | $1,066 | flat |
| Claude Sonnet 4.6 | Anthropic | 1M | none | 3 | $1,230 | flat |
| Claude Sonnet 4.5 | Anthropic | 200K | none | 3 | $1,230 | flat |
| Claude Opus 4.6 | Anthropic | 1M | none | 5 | $2,050 | flat |
| Claude Opus 4.5 | Anthropic | 200K | none | 5 | $2,050 | flat |
| GPT-5.6 Sol | OpenAI | 1.1M | none | 5 | $2,070 | flat |
| GPT-5.5 | OpenAI | 272K | none | 5 | $2,070 | flat |
| Claude Opus 5 | Anthropic | 1M | none | 5 | $2,665 | flat |
| Claude Opus 4.8 | Anthropic | 1M | none | 5 | $2,665 | flat |
| Claude Opus 4.7 | Anthropic | 1M | none | 5 | $2,665 | flat |
| Claude Fable 5 | Anthropic | 1M | none | 10 | $5,330 | flat |
| Claude Mythos 5 | Anthropic | 1M | none | 10 | $5,330 | flat |
| GPT-5.5 Pro | OpenAI | 272K | none | 30 | $12,420 | flat |
Token counts are scaled per model where a provider documents a different tokenizer — Claude 4.7 and later produce roughly 30% more tokens for the same text, which brings them to their context ceiling sooner even though their pricing stays flat. Caching and batch discounts are off here so the tier effect is visible on its own; model the full picture in the pricing calculator.
Where every cliff sits
Two structures exist in the market and they are not labelled as such on any pricing page. xAI and Google tier their long-context pricing; Anthropic, OpenAI, Google, DeepSeek publish a single rate across the whole window for the models listed below.
| Model | Cliff at | Input $/M | Output $/M | Cached $/M |
|---|---|---|---|---|
| Grok Build 0.1 · xAI | 200K | 1 → 2 | 2 → 4 | 0.2 → 0.4 |
| Gemini 2.5 Pro · Google | 200K | 1.25 → 2.5 | 10 → 15 | 0.125 → 0.25 |
| Grok 4.3 · xAI | 200K | 1.25 → 2.5 | 2.5 → 5 | 0.2 → 0.4 |
| Gemini 3.1 Pro (preview) · Google | 200K | 2 → 4 | 12 → 18 | 0.2 → 0.4 |
| Grok 4.6 · xAI | 200K | 2 → 4 | 6 → 12 | 0.5 → 1 |
| Grok 4.5 · xAI | 200K | 2 → 4 | 6 → 12 | 0.3 → 0.6 |
Every other model in the dataset (39 of them) bills a single rate across its full context window. Flat is not the same as cheap — several tiered models are cheaper than the flat ones right up until the threshold.
Why this is worth checking
A tiered price looks harmless on a pricing page: two columns, one for short prompts and one for long ones. The consequence only appears in a bill. Because the tier is chosen by total prompt size and then applied to the whole request, cost as a function of prompt length is a step function. A retrieval pipeline that gradually grew its context from 180k to 210k tokens does not see costs rise by 15% — it sees them roughly double, including on the output tokens, which did not change at all.
Nothing in the API signals this. There is no warning field in the response, no header, no error. The request succeeds and the invoice arrives at the end of the month.
Frequently asked questions
Does a longer prompt cost more per token?
On some models, yes — and the jump is a step, not a slope. 6 of the models tracked here switch to a higher rate once the prompt passes 200,000 tokens, and that higher rate is applied to the entire request rather than only the tokens above the threshold. A prompt just under the threshold and one just over it can differ in price by a factor of 2.00. The remaining 39 models charge one rate across their whole context window.
Is the boundary token itself billed at the low rate or the high rate?
It depends on the provider, and they word it differently. xAI says a request whose prompt reaches the threshold is billed at the higher rate, which reads as ≥ 200,000 tokens. Google's tables are split as "≤ 200K" and "> 200K", which puts the boundary token itself on the low rate. The practical impact is one token in a 200,000-token request, so it will not change a budget — but if you are building billing reconciliation against a provider invoice, the two are not the same rule. This calculator places the boundary above the threshold for every model; treat the exact boundary token as provider-specific rather than assuming a shared convention.
Is the higher rate applied only to the tokens above the threshold?
No, and this is the part that surprises people. xAI documents it plainly: a request whose prompt reaches the threshold is billed at the higher rate for all tokens in that request. Google's Gemini tiers work the same way — the rate is selected by the total prompt size, then applied to the whole prompt. There is no blended or marginal rate. This is why the cost curve has a vertical step in it rather than a bend.
Which models have a long-context price tier?
Grok Build 0.1 (xAI), $1 → $2 input and $2 → $4 output above 200,000 tokens; Gemini 2.5 Pro (Google), $1.25 → $2.5 input and $10 → $15 output above 200,000 tokens; Grok 4.3 (xAI), $1.25 → $2.5 input and $2.5 → $5 output above 200,000 tokens; Gemini 3.1 Pro (preview) (Google), $2 → $4 input and $12 → $18 output above 200,000 tokens; Grok 4.6 (xAI), $2 → $4 input and $6 → $12 output above 200,000 tokens; Grok 4.5 (xAI), $2 → $4 input and $6 → $12 output above 200,000 tokens. Every one of these thresholds sits at 200,000 tokens.
Does the long-context tier raise output prices too?
Yes, and it is easy to miss because output length has nothing to do with it. On the tiered models, a long prompt moves output onto the higher rate as well — the same 2,000-token answer costs more purely because the question was long. xAI doubles both input and output. Google raises input by 2× but output by 1.5× on Gemini 2.5 Pro and Gemini 3.1 Pro, so the two providers are not interchangeable even where the headline input prices match.
Does Claude charge more for its 1M-token context window?
No. Anthropic states that Claude 4.6 and later models include the full 1M token context window at standard pricing, and that a 900k-token request is billed at the same per-token rate as a 9k-token request. Prompt caching and batch discounts also apply at standard rates across the full window. That makes Claude's pricing flat, not cheap — the per-token rates are higher than several tiered models, so the flat structure matters most for workloads that genuinely live above 200k tokens.
How do I avoid crossing the cliff?
Measure first: the cliff only matters if your prompts actually approach it. If they do, the options in order of usual payoff are to trim retrieved context so the prompt stays under the threshold, split one long request into two shorter ones where the task allows it, or move that workload to a flat-rate model. Trimming is usually the biggest lever, because dropping below the threshold halves the price of the entire request — including the output tokens you did not change.
Does caching help once I am over the threshold?
It helps, but from a higher base. Cached input rates are tiered alongside the standard rates, so a cache read above the threshold costs more than a cache read below it. Caching and staying under the threshold are separate savings that multiply rather than substitute for each other.
Sources
- xAI — official pricing documentation
- Google — official pricing documentation
- Anthropic — official pricing documentation
- OpenAI — official pricing documentation
- DeepSeek — official pricing documentation
Related: full cost calculator · prompt cache analyzer · all model prices