Muse Spark 1.3 (contributor)
Frontier★ Best ValueMeta · $0.1/M in · $0.2/M out · $0.002/M cached
Sept 2026 — 10-20x cheaper, but Meta may train on your prompts and outputs
Estimate your monthly cost across all major AI APIs. Adjust the sliders, see what each model would cost.
32 models · prices verified May 2026 (base set) and September 23, 2026 (September launches)
Meta · $0.1/M in · $0.2/M out · $0.002/M cached
Sept 2026 — 10-20x cheaper, but Meta may train on your prompts and outputs
Google · $0.1/M in · $0.4/M out
OpenAI · $0.1/M in · $0.5/M out · $0.01/M cached
Sept 2026 — half GPT-5.6 Luna; cache writes $0.125; 128K output
OpenAI · $0.2/M in · $1.2/M out
Cut 80% from $1/$6 on 30 Jul 2026; succeeded by GPT-6 Luna at $0.10/$0.50
DeepSeek · $0.27/M in · $1.1/M out
Best price/perf ratio in frontier tier
OpenAI · $0.3/M in · $1.2/M out · $0.03/M cached
Mistral · $0.4/M in · $2/M out
Meta (via Together) · $0.88/M in · $0.88/M out
Google · $0.3/M in · $2.5/M out · $0.08/M cached
xAI · $1.25/M in · $2.5/M out · $0.2/M cached
Apr 2026 — the cheaper 1M-context option; 200K+ prompts bill at 2x
OpenAI · $1.1/M in · $4.4/M out
Affordable reasoning model
Meta · $1.25/M in · $4.25/M out · $0.15/M cached
Sept 2026 — standard endpoint; tops DeepSWE
Anthropic · $1/M in · $5/M out · $0.1/M cached
Mistral · $2/M in · $6/M out
xAI · $2/M in · $6/M out · $0.5/M cached
Sept 2026 — same price as Grok 4.6; prompts of 200K+ tokens bill $4/$12 for the whole request
xAI · $2/M in · $6/M out · $0.5/M cached
Aug 2026 — succeeded by Grok 4.7 at the same price
Google · $1.25/M in · $10/M out · $0.31/M cached
1M token context
Anthropic · $2/M in · $10/M out · $0.2/M cached
Jun 2026 — $2/$10 launch price made permanent; the planned 1 Sept rise to $3/$15 was cancelled
OpenAI · $2/M in · $10/M out · $0.1/M cached
29 Sept 2026 (DevDay) — GPT-6 Sol's rates with cached input halved; cache writes $2.50; prompts over 272K bill $4/$15
OpenAI · $2/M in · $10/M out · $0.2/M cached
22 Sept 2026 — half GPT-5.6 Sol's promotional $4/$20; cache writes $2.50; 128K output; succeeded by GPT-6.1 Sol on 29 Sept
OpenAI · $2/M in · $12/M out
Cut 20% from $2.50/$15 on 30 Jul 2026
Meta (via Together) · $5/M in · $5/M out
Open weights — also runnable locally for free
Anthropic · $3/M in · $15/M out · $0.3/M cached
Anthropic · $4/M in · $20/M out · $0.2/M cached
Sept 2026 — 20% under Opus 5; cache reads $0.20 (5% of input); 128K output; thinking always on
OpenAI · $4/M in · $20/M out · $0.4/M cached
Jun 2026 at $5/$30 — promotional $4/$20 since 21 Aug (through at least 21 Nov); succeeded by GPT-6 Sol at $2/$10
Anthropic · $5/M in · $25/M out · $0.5/M cached
Jul 2026 — now listed as legacy; succeeded by Opus 5.5 at $4/$20
Anthropic · $5/M in · $25/M out · $0.5/M cached
New tokenizer may add 0-35% to effective cost vs 4.6
OpenAI · $5/M in · $30/M out · $0.5/M cached
2x price increase vs GPT-5.4; Batch mode 50% off
OpenAI · $10/M in · $40/M out
Reasoning model — slower, much more capable on math/logic
Anthropic · $10/M in · $50/M out · $0.25/M cached
Sept 2026 — cache reads cut 75% vs Fable 5; 128K max output
OpenAI · $10/M in · $50/M out · $1/M cached
Sept 2026 — prompts over 272K tokens bill at 2x input; cache writes $12.50
OpenAI · $30/M in · $180/M out
Higher reasoning — Pro/Business/Enterprise
Input tokens = everything the model reads (your prompt, attached files, conversation history). Output tokens = what the model writes back.
Rule of thumb: 1 token ≈ 0.75 words. A 1,000-word document is ~1,300 tokens.
Cached input ratio matters because most providers offer steep discounts on tokens that hit their cache — Anthropic reads cache at a fortieth of the fresh-input price on Fable 5.1, OpenAI at a tenth on GPT-6 Astra. For applications with consistent system prompts or repeated context, cache hit rates of 50–90% are achievable.
The list is sorted cheapest first and the top card is marked best value, but that is best value for the exact tokens you entered. The calculator cannot know how many tokens a model will actually spend on your task, and that is where the real gap between models lies. Our GPT-6 Astra reviewis the case in point: Astra and Claude Fable 5.1 carry the identical $10/$50 list price and will show identical bills here, yet on Artificial Analysis's measured runs Fable 5.1 spent 2.7× more to finish the same benchmark because it emitted three times the output tokens. When two models are close on this page, use the cost-per-task column on the leaderboardto break the tie, and treat reasoning-heavy “max effort” settings as an output multiplier you are choosing to pay for.
Read the tier badges as a hint about behaviour, not quality. Frontier models write longer, think longer and cost more per token; Fast & Cheapmodels are the right default for high-volume classification, extraction and routing where a frontier model's judgement is wasted. Reasoningmodels bill their hidden thinking as output tokens, so a small visible answer can carry a large output bill. The notes under each card record repricings — GPT-5.6 Sol went to a promotional $4/$20 on 21 August and GPT-6 Sol then halved it to $2/$10, Sonnet 5's $2/$10 launch price was made permanent instead of rising to $3/$15, Claude Opus 5.5 arrived a fifth under Opus 5 at $4/$20 — because the date a price changed is as useful as the price.
Every figure is a provider's published API list price in US dollars per million tokens. The base table was verified against the official pricing pages in May 2026; the September 2026 frontier launches were added from published launch pricing, last on September 23, 2026. Sources: Anthropic, OpenAI, Google, xAI; other providers as linked from their tool pages. Providers update prices regularly — check the official page before committing to a production budget. Open-weight models like Llama 4 405B and DeepSeek V4 can also be self-hosted for unlimited usage if you have the hardware, which this calculator does not model.
What is not included: batch-API discounts (typically 50% in both directions), long-context surcharges (GPT-6 Astra bills prompts above 272K tokens at 2× input and 1.5× output), cache write fees where a provider charges them, image or audio tokens, and subscription plans. For subscriptions the question is different — whether a $20 seat pays for itself — and our Claude and ChatGPT Plus reviews answer it; for metered coding tools, the AI-credit explainer shows how a credit balance converts to tokens at these same rates.
For each model: (uncached input tokens × input price + cached input tokens × cached price + output tokens × output price) per request, multiplied by requests per day and by 30. Prices are the providers' published rates per million tokens; cached input uses the provider's cache-read price where one is published, otherwise the normal input price.
Only for identical output. The calculator prices the tokens you tell it about; a weaker model that needs three attempts, or a verbose one that writes three times the output, costs more in practice. For a like-for-like measure, use the cost-per-task column on the leaderboard, which records what each model actually spent to complete the same benchmark.
The share of your input tokens that hit the provider's prompt cache — a repeated system prompt, shared documents, or conversation history. Cache reads are billed at a steep discount (Claude Fable 5.1 reads cache at $0.25 per million against $10 for fresh input), so a stable prefix can cut input cost by most of its value.
The base table was verified against the providers' pricing pages in May 2026; the September 2026 launches (Claude Opus 5.5, Claude Fable 5.1, GPT-6 Astra, Sol and Luna, Grok 4.7, Muse Spark 1.3) were added from published launch pricing, most recently on September 23, 2026. Providers change prices without notice, so confirm on the official page before committing a budget.
No. It models standard synchronous pricing only. Batch APIs typically halve both directions, and some models charge more above a context threshold — GPT-6 Astra bills prompts over 272K tokens at 2× input — so treat the figure as a baseline, not an invoice.