DeepSeek Gained Ten Index Points Without Changing the Model — and Undercut Luna a Day After Its 80% Cut
DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, up from 40 in April, with identical architecture, the same 13B active parameters and the same price list — the gain came entirely from post-training. It now sits one point behind GPT-5.6 Luna at roughly 60% lower cost per task, with a 98% cache discount. The nuance most coverage will skip: accuracy was unchanged, and the jump came mainly from a reduced hallucination rate.
DeepSeek shipped a point update to its budget model today, and the numbers are strange in an interesting way. DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index — a ten-point jump over the V4 Flash released in April, which scored 40. It now sits one point behind GPT-5.6 Luna, at roughly 60% lower cost per task.
The strange part: nothing about the model got bigger. Artificial Analysis reports it "shares identical architecture and pricing with the earlier DeepSeek V4 Flash" — same 13B active parameters out of 284B total, same price list. The entire gain came from post-training.
It also arrives one day after OpenAI cut Luna's price by 80%. That cut is already the second-best deal in this comparison.
Where it lands
| Model | Intelligence Index | Price per 1M (in / out) |
|---|---|---|
| GPT-5.6 Luna (max) | 51 | $0.20 / $1.20 |
| DeepSeek V4 Flash 0731 | 50 | $0.14 / $0.28 |
| DeepSeek V4 Pro | 44 | — |
| DeepSeek V4 Flash (April) | 40 | same as 0731 |
Two things stand out. The small model now beats DeepSeek's own Pro tier by six points, which is an awkward result for the company's own lineup. And on output tokens — where the real money goes in generation-heavy work — it is roughly a quarter of Luna's post-cut price.
Ten points, and almost none of it from being more right
This is the detail that most coverage of the release will skip, and it changes what the number means.
Artificial Analysis attributes the improvement primarily to a reduced hallucination rate, with overall accuracy — percentage of answers correct — unchanged. The model did not get better at solving problems. It got substantially better at not confidently asserting things that are false.
Read pessimistically, that deflates the headline: "ten points better" sounds like a capability leap and is not one. Read practically, it may be the more useful improvement. In an agent loop, a confidently wrong intermediate step poisons everything downstream, and a model that declines to invent an answer is worth more than one that is right marginally more often. DeepSeek also reports the model using 12% fewer output tokens than its predecessor, which is consistent with less rambling and less confabulation.
It is a genuinely notable engineering result either way: a ten-point index move extracted from weights that already existed, with no architecture change and no additional parameters.
The cache discount is the actual weapon
The headline price is $0.14 per million input tokens and $0.28 per million output. The number that matters more for anyone running agents is the cached input price: $0.0028 per million tokens, a 98% cache discount against an industry norm closer to 90%.
That difference sounds marginal and is not. If your workload re-sends a large stable prefix on every call — a long system prompt, a codebase, a document set — the cached portion dominates your input bill, and the gap between a 90% and a 98% discount is a 5× difference on that portion. This is priced for exactly the workload the industry is moving toward.
What it does to yesterday's price cut
We wrote yesterday that OpenAI's 80% cut on Luna was a competitive move rather than an efficiency pass-through, and that the entry tier was being defended as a commodity. That thesis held up faster than expected — within twenty-four hours, the freshly cut price was undercut by roughly 60% on cost per task by a model one index point behind.
One correction to our own framing while we are here: that piece said the thing to watch was whether Anthropic and Google would follow at the bottom of their ranges. That was the wrong place to look. The pressure came from a Chinese lab, it took a day rather than a quarter, and it arrived through better post-training rather than a price cut at all — DeepSeek did not change its prices. It changed what you get for them.
This is the pattern we described back in June, when frontier prices started moving up while the bottom of the market raced toward zero. The bifurcation is now sharp enough to see week by week.
Context from our leaderboard
An index of 50 is not a frontier score. On our model leaderboard — captured on July 25, so it does not yet include this release — 50 would sit in the same band as Gemini 3.5 Flash (50.2) and Meta's Muse Spark 1.1 (50.6), roughly fifteenth, well below Claude Opus 5 at 61 and GPT-5.6 Sol at 58.9.
That is the correct way to hold this story. DeepSeek has not built a frontier model. It has built something mid-table on capability that costs a fraction of anything nearby, which is a different and in some ways more commercially disruptive achievement. It is the same playbook that produced V4 in April and the V4 preview before it.
The counter-case
A one-point gap is noise. The Intelligence Index is a composite of roughly ten benchmarks. Treating 50 versus 51 as a real ranking is over-reading a single aggregate number, and cost-per-task figures additionally depend on token-efficiency assumptions that vary by workload.
The cache advantage is conditional. A 98% discount is transformative for long stable prefixes and nearly irrelevant for short, varied prompts. Plenty of production traffic looks like the latter.
Accuracy did not move. If your bottleneck is the model getting hard problems right rather than being calibrated about uncertainty, this update does not help you.
And there is the provider question. Running production traffic through a Chinese-hosted API raises data-residency and procurement questions that many Western enterprises will not clear, and they are more live today than usual given this morning's Reuters review of Chinese military distillation of US models. Earlier V4 models were released as open weights, which would make self-hosting an answer — but we could not confirm that weights for the 0731 update have been published, so treat that as unresolved rather than assumed.
Worth noting too that these figures come from Artificial Analysis's independent testing rather than DeepSeek's own benchmark claims, which is the stronger form of evidence — and the same index that placed Kimi K3 near the frontier when it launched.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.