Meta's Muse Spark 1.3 Leads Coding at $0.10 a Million
Meta's fourth frontier model in five months takes the top spot on DeepSWE, near-saturates million-token retrieval, and undercuts every rival on price — as long as you let Meta train on your traffic.
Meta shipped Muse Spark 1.3 on 2 September, roughly four weeks after 1.2 and the fourth frontier release out of Meta Superintelligence Labs since April. It is live now inside the Muse Code coding agent and on the Meta Model API, and the headline claim is narrow but real: on coding benchmarks, Meta is no longer chasing the field.
The clearest jump is on DeepSWE v1.1, where 1.3 scores 75.4 — ahead of Claude Opus 5 at 74.0 and GPT-5.6 Sol at 73.0, and a long way above the 55.0 that Muse Spark 1.2 managed four weeks earlier. It also leads SWEAtlas CodeBase QnA at 59.4 against 53.5 for GPT-5.6 Sol and 52.7 for Opus 5, and ties GPT-5.6 Sol on Terminal-Bench 2.1 at 88.8. Long-context retrieval is close to saturated: 98.5 on MRCR at 256K to 512K and 98.1 at 512K to 1M, where GPT-5.6 Sol falls to 91.5 and 73.8 respectively. The model takes text, image and video input across a 1M-token window.
Agentic work is where the picture flattens out. On Artificial Analysis numbers the xhigh variant lands at Intelligence Index 61 — sixth of 636 models tracked, at about 182 tokens per second — and Opus 5 in max mode still edges it on every agent row: GDPVal-AA v2 1,824 to 1,754, JobBench 65.7 to 64.9, OSWorld 2.0 68.3 to 66.9, DeepSearchQA 90.4 to 89.4. Meta’s own max reasoning tier, which might close that gap, was still gated behind safety testing at launch. What Meta does claim across the board is efficiency: in internal engineer comparisons, 1.3 finishes coding tasks with roughly 20% fewer tool calls and 25% fewer tokens than 1.2, which matters more than a benchmark point when an agent has been running for an hour.
The genuinely unusual part is the price sheet. The standard xhigh endpoint runs $1.25 per million input tokens and $4.25 per million output, with cached input at $0.15 — an 88% cache discount, and roughly in line with the rest of the frontier. Sitting beside it is a “contributor” endpoint at about $0.10 in and $0.20 out, with cached input at $0.002. That is a 10x to 20x discount, and the consideration is your data: prompts and outputs on the contributor tier may be used to improve Meta’s products. Rivals sell privacy as the default and charge for it. Meta has instead put a number on what that default is worth, and invited developers to sell it back.
Mark Zuckerberg framed the release on X as frontier performance almost too cheap to meter, and used the same post to promise two things that were not actually in the launch: open weights for the Spark line, and a release codenamed only with a watermelon emoji, both “coming soon.” The open-weights promise has some history behind it — Meta opened the weights for the 30B Muse Glimmer in August, and a similar “soon” was attached to Muse Spark 1.2 at the time. As of 1.3, the flagship remains a closed API model.
Rollout is uneven too. Users in the EU reported that Meta.ai was still serving Muse Spark 1.1 after 1.3 went live on the API, so the consumer surface and the developer surface are not moving together. For anyone already paying for Muse Code at $5 a month the upgrade arrives with no extra step. And for anyone benchmarking coding agents this week, the honest read is that four labs now sit within a couple of points of each other on the tests that get quoted, which leaves price and token efficiency doing more of the deciding than raw capability is.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.