Models·3 min read
By BitsMindsSource: OpenRouter

DeepSeek Shipped V4-Pro Without Telling Anyone

No blog post, no changelog, no press release — just an edit to the API pricing page. DeepSeek-V4-Pro-0813 is now generally available at $0.87 per million output tokens, with benchmark gains that nobody outside DeepSeek has reproduced.

$0.87 / 1M OUTPUT V4-PRO-0813 GENERALLY AVAILABLE BITSMINDS.COM
Share:

DeepSeek moved its flagship model to general availability on 12 August by editing a pricing table. There was no blog post, no changelog entry and no press release. The deepseek-v4-pro endpoint now resolves to a build tagged V4-Pro-0813, ending a preview period that had run since 24 April — close to four months, and long enough that a fair number of people had started treating the preview as the product.

The commercial terms are the part that is not in dispute. OpenRouter lists the model at $0.435 per million input tokens and $0.87 per million output tokens, with cache hits at $0.003625 per million and a 1M-token context window on a large mixture-of-experts architecture. DeepSeek has also said an API price increase is coming "in the near future" — with no date, no figure, and no indication of which models or tiers it will touch. Today's numbers are explicitly temporary.

The performance claims are a different matter. Against the April preview, DeepSeek reports DeepSWE rising from 12.8 to 62.7, CyberGym from 52.7 to 83.3, and Terminal Bench 2.1 from 72.1 to 87.9. A near-50-point jump on an agentic coding benchmark inside one model generation is not impossible, but it is the kind of result that normally arrives with a technical report attached. This one did not. No third party has reproduced any of it, and benchable.ai's page for the model currently shows an empty verification section.

That gap matters more than usual because of how the numbers are being used. The line circulating this week is that V4-Pro is roughly a fifty-seventh of the price of Fable 5 at comparable capability. That ratio holds only if you compare output tokens in isolation; on a realistic blended input-output workload it lands closer to a forty-sixth. Still an enormous discount — but the headline figure is the flattering slice of a real one, and the capability half of the comparison rests entirely on self-reported scores.

The open-weights question is also unresolved. Earlier V4 builds went out with weights on Hugging Face, which is a large part of why DeepSeek's releases get scrutinised so quickly by outside researchers. The 0813 build's weights have not been published. Until they are, the only way to evaluate this model is to pay for tokens and test it yourself, which is precisely the situation open weights were supposed to avoid.

None of this is out of character. DeepSeek has spent 2026 shipping on its own cadence and letting the price sheet do the announcing — V4-Flash landed the same way at the end of July, and the April preview was similarly light on ceremony. For a lab whose entire competitive position is cost per token rather than launch-day narrative, that is arguably rational. It also means the burden of verification has been quietly transferred to everyone else.

The practical read for anyone deciding whether to route traffic here: the API is live, the context window is real, and the price is real for now. The benchmark deltas are a vendor's claim about its own product, and they should be treated as one until an independent harness says otherwise. Given the four-month preview, the first credible external numbers should not be far off.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Four models, three miniature worlds An isometric cloverleaf interchange with a raised bridge and tiny cars, a seaside Ferris wheel and carousel, and a rocket ascending toward a satellite form three detailed model-making dioramas. The header names Claude Sonnet 5.5, Claude Opus 5.5, GPT-6 Astra and GPT-6 Sol. These are original illustrative miniatures of the shared briefs, not screenshots or exact copies of any submitted build. Their sizes, positions and colours do not encode scores or a ranking. BITSMINDS LAB One attempt. Three builds. CLAUDESonnet 5.5VSCLAUDEOpus 5.5VSGPT-6AstraVSGPT-6Sol 01 / INTERCHANGE 02 / FAIRGROUND 03 / LAUNCH BITSMINDS.COM
Models

Claude Sonnet 5.5 vs Opus 5.5, Astra, Sol: A Point Short

Claude Sonnet 5.5: more capability at Sonnet prices An ivory, machined performance console has a large terracotta rotary dial bearing the official Claude asterisk. Its pointer is set to High among five effort settings: Low, Medium, High, Xhigh and Max. Beside it, the title names Claude Sonnet 5.5 and quotes its standard API prices of 2 US dollars per million input tokens and 10 US dollars per million output tokens. The physical console is an editorial metaphor, not an Anthropic product. The highlighted effort setting reflects the linked article's independent cost/performance discussion and does not promise an optimal setting for every workload. LOWMEDIUMHIGHXHIGHMAX EFFORT HIGH LOW MEDIUM SONNET 5.5 ADAPTIVE THINKING CLAUDE Sonnet 5.5 $2 INPUT $10 OUTPUT USD PER MILLION TOKENS BITSMINDS.COM
Models

Claude Sonnet 5.5 Is Out: Sonnet Price, Near-Opus Scores

Claude Sonnet 5.5: a model on the horizon, details still open The official Claude asterisk becomes a floating terracotta sculpture above a cream circular plinth. A calendar and a price tag orbit it, each carrying a question mark. The title names Sonnet 5.5; the accompanying caption makes clear that its release date and price are unconfirmed. The artwork illustrates the linked article's distinction between a forthcoming model and uncertain launch details, without asserting a date, price or specification. COMING WEEKS RELEASE API PRICE $ ? ANTHROPIC Sonnet 5.5 Release date & price STILL UNCONFIRMED BITSMINDS.COM
Models

Claude Sonnet 5.5: What the Opus 5.5 Leaker Says Now