LLM API Pricing Calculator

Estimate what an LLM workload actually costs per month across 45 models from 5 providers — Anthropic, OpenAI, Google, xAI and DeepSeek. Unlike most calculators, this one accounts for prompt caching, batch discounts, long-context price tiers, and the fact that different models tokenize the same text into different numbers of tokens.

Prices in USD per million tokens · verified 2026-08-16 · runs entirely in your browser

Start from a workload
Input tokens / request
Output tokens / request
Requests / month
Cache hit rate — 60%
Why your token counts differ per model. Claude 4.7 and later (including Opus 5, Sonnet 5, Fable 5) use a newer tokenizer that produces roughly 30% more tokens for the same text. A lower sticker price on a model that tokenizes your text less efficiently can still cost more. This calculator scales token counts per model so the comparison reflects real spend — most calculators skip this.
Providers
Cheapest for this workload
GPT-5 nano$27.80/month
That is $24,272/month less than GPT-5.5 Pro (99.9% cheaper).
ModelProvider$/M in$/M outBilled tokens/reqCost/requestMonthlyvs cheapest
GPT-5 nanoOpenAI0.050.412.7K$0.000556$27.80baseline
Gemini 2.5 Flash-LiteGoogle0.10.412.7K$0.000832$41.601.5×
DeepSeek V4 FlashDeepSeek0.140.2812.7K$0.000888$44.411.6×
GPT-4o miniOpenAI0.150.612.7K$0.00168$84.003.0×
GPT-5.6 LunaOpenAI0.21.212.7K$0.00194$97.203.5×
GPT-5.4 nanoOpenAI0.21.2512.7K$0.00198$98.953.6×
Gemini 3.1 Flash-LiteGoogle0.251.512.7K$0.00243$121.504.4×
DeepSeek V4 ProDeepSeek0.4350.8712.7K$0.00272$136.164.9×
GPT-5 miniOpenAI0.25212.7K$0.00278$139.005.0×
Gemini 3.5 Flash-LiteGoogle0.32.512.7K$0.00341$170.306.1×
Gemini 2.5 FlashGoogle0.32.512.7K$0.00341$170.306.1×
GPT-4.1 miniOpenAI0.41.612.7K$0.00376$188.006.8×
Gemini 3.7 FlashGoogle0.753.7512.7K$0.00677$338.2512.2×
Gemini 3.6 FlashGoogle0.753.7512.7K$0.00677$338.2512.2×
GPT-5.4 miniOpenAI0.754.512.7K$0.00729$364.5013.1×
Grok Build 0.1xAI1212.7K$0.00764$382.0013.7×
Claude Haiku 4.5Anthropic1512.7K$0.00902$451.0016.2×
Grok 4.3xAI1.252.512.7K$0.00919$459.5016.5×
o4-miniOpenAI1.14.412.7K$0.0103$517.0018.6×
GPT-5.1OpenAI1.251012.7K$0.0139$695.0025.0×
GPT-5OpenAI1.251012.7K$0.0139$695.0025.0×
Gemini 3.5 FlashGoogle1.5912.7K$0.0146$729.0026.2×
Grok 4.5xAI2612.7K$0.016$798.0028.7×
Grok 4.6xAI2612.7K$0.0174$870.0031.3×
GPT-4.1OpenAI2812.7K$0.0188$940.0033.8×
o3OpenAI2812.7K$0.0188$940.0033.8×
GPT-5.6 TerraOpenAI21212.7K$0.0194$972.0035.0×
Gemini 3.1 Pro (preview)previewGoogle21212.7K$0.0194$972.0035.0×
GPT-5.3 CodexOpenAI1.751412.7K$0.0195$973.0035.0×
GPT-5.2OpenAI1.751412.7K$0.0195$973.0035.0×
Gemini 2.5 ProGoogle1.251012.7K$0.022$1,10039.6×
Claude Sonnet 5Anthropic21016.5K ×1.3$0.0235$1,17342.2×
GPT-5.4OpenAI2.51512.7K$0.0243$1,21543.7×
Claude Sonnet 4.6Anthropic31512.7K$0.0271$1,35348.7×
Claude Sonnet 4.5Anthropic31512.7K$0.0271$1,35348.7×
Claude Opus 4.6Anthropic52512.7K$0.0451$2,25581.1×
Claude Opus 4.5Anthropic52512.7K$0.0451$2,25581.1×
GPT-5.6 SolOpenAI53012.7K$0.0486$2,43087.4×
GPT-5.5OpenAI53012.7K$0.0486$2,43087.4×
Claude Opus 5Anthropic52516.5K ×1.3$0.0586$2,931105.4×
Claude Opus 4.8Anthropic52516.5K ×1.3$0.0586$2,931105.4×
Claude Opus 4.7Anthropic52516.5K ×1.3$0.0586$2,931105.4×
Claude Fable 5Anthropic105016.5K ×1.3$0.117$5,863210.9×
Claude Mythos 5limitedAnthropic105016.5K ×1.3$0.117$5,863210.9×
GPT-5.5 ProOpenAI3018012.7K$0.486$24,300874.1×

The four things that actually move your bill

Teams usually compare the headline $/1M tokens figure and stop there. In practice that number explains less than half of a real invoice. These four factors explain most of the rest.

1. Tokenizer efficiency

Providers do not charge for characters, they charge for tokens — and every model family splits text differently. Anthropic states that Claude 4.7 and later produce about 30% more tokens for the same text than earlier Claude models. A model priced 20% lower but tokenizing 30% less efficiently is more expensive, not less. Toggle Correct for tokenizer differences above to see the effect.

2. Prompt caching

Cache reads cost roughly 10% of base input price across Anthropic, OpenAI and Google. Any workload that re-sends a stable prefix — a system prompt, a tool schema, a retrieved corpus — should be caching. For agent workloads with 80% cache hit rates this is frequently a larger saving than switching model families.

3. Batch processing

A flat 50% discount on Anthropic and OpenAI for asynchronous work. If a job can tolerate latency, not using batch is leaving half the money on the table.

4. Long-context tiers

Google and xAI roughly double their rates above 200k input tokens. Anthropic prices its full 1M-token window flat. If your inputs straddle 200k, this single difference can outweigh every other factor.

Frequently asked questions

Which LLM API is cheapest in 2026?

On sticker price alone, the cheapest general-purpose models are GPT-5 nano ($0.05 per million input tokens), Gemini 2.5 Flash-Lite ($0.10) and DeepSeek V4 Flash ($0.14). But sticker price rarely decides the bill: prompt caching, batch discounts, long-context price tiers and tokenizer efficiency routinely change the ranking. For a cached, long-context workload a mid-tier model can beat a "cheaper" one outright.

Why do two models with the same price per token cost different amounts?

Because a token is not a fixed amount of text. Claude 4.7 and later models (Opus 5, Sonnet 5, Fable 5) use a newer tokenizer that produces roughly 30% more tokens for the same input than Claude Sonnet 4.6 and earlier. Two providers quoting the same $/million rate can therefore differ by ~30% on an identical prompt. This calculator applies a per-model tokenizer factor so the comparison reflects the actual bill.

How much does prompt caching actually save?

On Anthropic models a cache hit costs 10% of the base input price, while writing to the cache costs 1.25× (5-minute TTL) or 2× (1-hour TTL). So a 5-minute cache pays for itself after a single re-read, and a 1-hour cache after two. OpenAI and Google cached input is typically 10% of base input as well. DeepSeek is the outlier: since April 2026 a V4 Flash cache hit costs $0.0028 per million tokens against $0.14 for a miss — a 50× discount. For agent and RAG workloads, where a large system prompt or document set is re-sent every turn, caching is usually the single largest lever on cost, often larger than switching models.

When is the Batch API worth using?

Anthropic and OpenAI both discount batch requests by 50% on input and output. Batch is asynchronous, so it fits classification, extraction, evaluation and backfill jobs, but not interactive traffic. Note that batch does not stack with every feature — Anthropic's fast mode, for instance, is unavailable in batch. xAI and DeepSeek do not currently publish a batch discount.

Do long inputs cost more per token?

On some models, yes. Gemini 3.1 Pro and Gemini 2.5 Pro roughly double input pricing above 200k tokens, and Grok models double both input and output above 200k. Anthropic takes the opposite approach: Claude 4.6 and later include the full 1M-token context window at standard pricing, so a 900k-token request is billed at the same per-token rate as a 9k-token one. This calculator applies the correct tier automatically.

Where does this pricing data come from?

Every number is taken from the provider's official pricing documentation and re-verified on a schedule. The dataset was last verified on 2026-08-16. Where a provider's published figure is ambiguous, the field is left blank rather than guessed — you will see "—" instead of a number we are not confident in.

Sources