Long-Context Price Cliff Calculator

6 of the models tracked here switch to a higher rate once a prompt passes 200,000 tokens. The higher rate is not charged on the overflow — it is charged on every token in the request, output included. On Grok 4.6, a prompt just under 200,000 tokens and one just over it differ in price by 2.00×. One token.

Runs in your browser · tier pricing verified 2026-08-18 against official provider docs

Prompt size
At 195K input tokens
No model is on its long-context rate yet
Cheapest here: GPT-5 nano at $21.10/month.
You are 5,000 tokens away from a price cliff. Cross it and Gemini 3.1 Pro (preview) goes from $828.00 to $1,672 per month — an increase of $844.01 for one extra token. The higher rate is not applied to the overflow; it is applied to every token in the request, output included.
ModelProviderContext windowCliff at$/M in (applied)Monthly hereMonthly past the cliff
GPT-5 nanoOpenAI—none0.05$21.10flat
Gemini 2.5 Flash-LiteGoogle—none0.1$40.60flat
GPT-4o miniOpenAI—none0.15$60.90flat
GPT-5.6 LunaOpenAI1.1Mnone0.2$82.80flat
GPT-5.4 nanoOpenAI—none0.2$83.00flat
Gemini 3.1 Flash-LiteGoogle—none0.25$103.50flat
GPT-5 miniOpenAI—none0.25$105.50flat
Gemini 3.5 Flash-LiteGoogle—none0.3$127.00flat
Gemini 2.5 FlashGoogle1Mnone0.3$127.00flat
GPT-4.1 miniOpenAI—none0.4$162.40flat
DeepSeek V4 FlashDeepSeek1Mnone0.44$176.88flat
Gemini 3.7 FlashGoogle—none0.75$307.50flat
Gemini 3.6 FlashGoogle—none0.75$307.50flat
GPT-5.4 miniOpenAI—none0.75$310.50flat
Grok Build 0.1xAI256K200K1$398.00$816.00 ×2.05
Claude Haiku 4.5Anthropic200Knone1$410.00flat
o4-miniOpenAI—none1.1$446.60flat
Grok 4.3xAI1M200K1.25$497.50$1,020 ×2.05
GPT-5.1OpenAI—none1.25$527.50flat
GPT-5OpenAI—none1.25$527.50flat
Gemini 2.5 ProGoogle—200K1.25$527.50$1,060 ×2.01
DeepSeek V4 ProDeepSeek1Mnone1.32$530.64flat
Gemini 3.5 FlashGoogle—none1.5$621.00flat
GPT-5.3 CodexOpenAI—none1.75$738.50flat
GPT-5.2OpenAI—none1.75$738.50flat
Grok 4.6xAI500K200K2$804.00$1,648 ×2.05
Grok 4.5xAI500K200K2$804.00$1,648 ×2.05
GPT-4.1OpenAI—none2$812.00flat
o3OpenAI200Knone2$812.00flat
GPT-5.6 TerraOpenAI1.1Mnone2$828.00flat
Gemini 3.1 Pro (preview)Google—200K2$828.00$1,672 ×2.02
GPT-5.4OpenAI272Knone2.5$1,035flat
Claude Sonnet 5Anthropic1Mnone2$1,066flat
Claude Sonnet 4.6Anthropic1Mnone3$1,230flat
Claude Sonnet 4.5Anthropic200Knone3$1,230flat
Claude Opus 4.6Anthropic1Mnone5$2,050flat
Claude Opus 4.5Anthropic200Knone5$2,050flat
GPT-5.6 SolOpenAI1.1Mnone5$2,070flat
GPT-5.5OpenAI272Knone5$2,070flat
Claude Opus 5Anthropic1Mnone5$2,665flat
Claude Opus 4.8Anthropic1Mnone5$2,665flat
Claude Opus 4.7Anthropic1Mnone5$2,665flat
Claude Fable 5Anthropic1Mnone10$5,330flat
Claude Mythos 5Anthropic1Mnone10$5,330flat
GPT-5.5 ProOpenAI272Knone30$12,420flat

Token counts are scaled per model where a provider documents a different tokenizer — Claude 4.7 and later produce roughly 30% more tokens for the same text, which brings them to their context ceiling sooner even though their pricing stays flat. Caching and batch discounts are off here so the tier effect is visible on its own; model the full picture in the pricing calculator.

Where every cliff sits

Two structures exist in the market and they are not labelled as such on any pricing page. xAI and Google tier their long-context pricing; Anthropic, OpenAI, Google, DeepSeek publish a single rate across the whole window for the models listed below.

Model Cliff at Input $/M Output $/M Cached $/M
Grok Build 0.1 · xAI 200K 1 → 2 2 → 4 0.2 → 0.4
Gemini 2.5 Pro · Google 200K 1.25 → 2.5 10 → 15 0.125 → 0.25
Grok 4.3 · xAI 200K 1.25 → 2.5 2.5 → 5 0.2 → 0.4
Gemini 3.1 Pro (preview) · Google 200K 2 → 4 12 → 18 0.2 → 0.4
Grok 4.6 · xAI 200K 2 → 4 6 → 12 0.5 → 1
Grok 4.5 · xAI 200K 2 → 4 6 → 12 0.3 → 0.6

Every other model in the dataset (39 of them) bills a single rate across its full context window. Flat is not the same as cheap — several tiered models are cheaper than the flat ones right up until the threshold.

Why this is worth checking

A tiered price looks harmless on a pricing page: two columns, one for short prompts and one for long ones. The consequence only appears in a bill. Because the tier is chosen by total prompt size and then applied to the whole request, cost as a function of prompt length is a step function. A retrieval pipeline that gradually grew its context from 180k to 210k tokens does not see costs rise by 15% — it sees them roughly double, including on the output tokens, which did not change at all.

Nothing in the API signals this. There is no warning field in the response, no header, no error. The request succeeds and the invoice arrives at the end of the month.

Frequently asked questions

Does a longer prompt cost more per token?

On some models, yes — and the jump is a step, not a slope. 6 of the models tracked here switch to a higher rate once the prompt passes 200,000 tokens, and that higher rate is applied to the entire request rather than only the tokens above the threshold. A prompt just under the threshold and one just over it can differ in price by a factor of 2.00. The remaining 39 models charge one rate across their whole context window.

Is the boundary token itself billed at the low rate or the high rate?

It depends on the provider, and they word it differently. xAI says a request whose prompt reaches the threshold is billed at the higher rate, which reads as ≥ 200,000 tokens. Google's tables are split as "≤ 200K" and "> 200K", which puts the boundary token itself on the low rate. The practical impact is one token in a 200,000-token request, so it will not change a budget — but if you are building billing reconciliation against a provider invoice, the two are not the same rule. This calculator places the boundary above the threshold for every model; treat the exact boundary token as provider-specific rather than assuming a shared convention.

Is the higher rate applied only to the tokens above the threshold?

No, and this is the part that surprises people. xAI documents it plainly: a request whose prompt reaches the threshold is billed at the higher rate for all tokens in that request. Google's Gemini tiers work the same way — the rate is selected by the total prompt size, then applied to the whole prompt. There is no blended or marginal rate. This is why the cost curve has a vertical step in it rather than a bend.

Which models have a long-context price tier?

Grok Build 0.1 (xAI), $1 → $2 input and $2 → $4 output above 200,000 tokens; Gemini 2.5 Pro (Google), $1.25 → $2.5 input and $10 → $15 output above 200,000 tokens; Grok 4.3 (xAI), $1.25 → $2.5 input and $2.5 → $5 output above 200,000 tokens; Gemini 3.1 Pro (preview) (Google), $2 → $4 input and $12 → $18 output above 200,000 tokens; Grok 4.6 (xAI), $2 → $4 input and $6 → $12 output above 200,000 tokens; Grok 4.5 (xAI), $2 → $4 input and $6 → $12 output above 200,000 tokens. Every one of these thresholds sits at 200,000 tokens.

Does the long-context tier raise output prices too?

Yes, and it is easy to miss because output length has nothing to do with it. On the tiered models, a long prompt moves output onto the higher rate as well — the same 2,000-token answer costs more purely because the question was long. xAI doubles both input and output. Google raises input by 2× but output by 1.5× on Gemini 2.5 Pro and Gemini 3.1 Pro, so the two providers are not interchangeable even where the headline input prices match.

Does Claude charge more for its 1M-token context window?

No. Anthropic states that Claude 4.6 and later models include the full 1M token context window at standard pricing, and that a 900k-token request is billed at the same per-token rate as a 9k-token request. Prompt caching and batch discounts also apply at standard rates across the full window. That makes Claude's pricing flat, not cheap — the per-token rates are higher than several tiered models, so the flat structure matters most for workloads that genuinely live above 200k tokens.

How do I avoid crossing the cliff?

Measure first: the cliff only matters if your prompts actually approach it. If they do, the options in order of usual payoff are to trim retrieved context so the prompt stays under the threshold, split one long request into two shorter ones where the task allows it, or move that workload to a flat-rate model. Trimming is usually the biggest lever, because dropping below the threshold halves the price of the entire request — including the output tokens you did not change.

Does caching help once I am over the threshold?

It helps, but from a higher base. Cached input rates are tiered alongside the standard rates, so a cache read above the threshold costs more than a cache read below it. Caching and staying under the threshold are separate savings that multiply rather than substitute for each other.

Sources

Related: full cost calculator · prompt cache analyzer · all model prices