Grok 4.6 vs Opus 5 vs GPT-5.6 Sol: Frontier, 60% Cheaper
Artificial Analysis scored Grok 4.6 at 61 on its Intelligence Index — level with GPT-5.6 Sol at max, behind Fable 5 at 62 and Opus 5 at 63. It costs $2/$6 per million tokens against $5/$25 and $5/$30, and finishes long agentic tasks in ~53 turns where Opus 5 takes ~103. It also trails both rivals by eight points on the hardest software-engineering evals.
SpaceXAI released Grok 4.6 on Wednesday, and Artificial Analysis scored it at 61 on its Intelligence Index — level with GPT-5.6 Sol at max effort, behind Claude Fable 5 at 62 and Claude Opus 5 at 63. That is a three-way frontier within two points, and the first time a Grok release has been measured inside it. The number that actually separates Grok 4.6 from the models it ties is on the invoice: $2 per million input tokens and $6 per million output, against $5/$25 for Opus 5 and $5/$30 for Sol.
The jump is the steepest in the model line's history. Grok 4.5 scored 56 on the same index roughly a month ago; Grok 4.3 was 23 points back. Our write-up of Grok 4.5 in July described a model that deliberately traded leaderboard position for efficiency — Opus-class output at a fraction of the tokens, without ever topping a chart. Grok 4.6 keeps the efficiency argument and stops conceding the chart. SpaceXAI attributes the gain to extended supplemental training on curated model-generated data, regenerated trajectories from Grok 4.5, and agentic reinforcement learning across software engineering and domain-specific environments.
The headline numbers
| Spec | Grok 4.6 | Claude Opus 5 | GPT-5.6 Sol |
|---|---|---|---|
| Intelligence Index | 61 (high) | 63 (max) | 61 (max) |
| Input / MTok | $2 | $5 | $5 |
| Output / MTok | $6 | $25 | $30 |
| Cached input / MTok | $0.50 | — | — |
| Context window | 500K | 1M | 1.05M |
The context window is unchanged from Grok 4.5 at 500,000 tokens, which leaves Grok the smallest window of the three — Opus 5 ships 1M. There is also a pricing cliff worth reading carefully before you budget: prompts at or above 200,000 tokens move to a long-context tier at $4/$12, and once a request crosses that threshold SpaceXAI bills every token in it at the long-context rate, not just the tokens past the line. A faster serving variant is offered at double the price, and the launch page does not spell out how the two multipliers combine.
Where it wins, and where it doesn't
SpaceXAI published a benchmark table against the top reasoning tiers of its rivals. The pattern in it is consistent enough to be useful:
| Benchmark | Grok 4.6 | Fable 5 Max | GPT-5.6 Sol Max |
|---|---|---|---|
| GDPval-AA v2 (Elo) | 1753 | 1741 | 1728 |
| AA-Briefcase (Elo) | 1577 | 1574 | 1502 |
| Harvey LAB | 15.8% | 11.3% | 2.5% |
| CursorBench v3.2 | 69.9% | 70.5% | 67.2% |
| FrontierCode v1.1 | 61.3% | 64.9% | 60.6% |
| APEX-Agents | 57.5% | 59.2% | 56.7% |
| APEX-SWE | 56.4% | 58.8% | — |
| DeepSWE v1.1 | 65.9% | 70.0% | 73.0% |
| Terminal-Bench v3.0 | 26.0% | 34.1% | 34.6% |
Grok 4.6 leads on the evals that measure economically-shaped work over long horizons — GDPval, AA-Briefcase, and the Harvey legal benchmark, where its 15.8% is more than six times Sol's 2.5%. It loses on the hard end of software engineering: eight points behind Fable 5 on DeepSWE, and only 26% on Terminal-Bench v3.0 against roughly 34% for both rivals. A model that can hold a long agentic task together but stumbles on gnarly terminal work is an unusual shape, and it is the opposite of what a raw index score of 61 implies about uniform strength.
One caution when reading coverage of this launch: Terminal-Bench appears at two different versions in circulation. SpaceXAI's table reports 26.0% on v3.0, while Artificial Analysis's own run reports 88.4% on v2.1. Those are different tests with different difficulty, not a contradiction, and they cannot be stacked into a single claim.
The efficiency argument, quantified
Price per token understates the gap, because Grok 4.6 also uses fewer of them. On long-horizon tasks Artificial Analysis measured it finishing in around 53 turns and 0.5 billion input tokens on average, against roughly 103 turns and 2.0 billion input tokens for Opus 5. Cost to run its index came to $0.84 per task, which Artificial Analysis puts at more than 60% below both Opus 5 and Sol. Fewer turns at a quarter the output rate compounds, and for an agent that runs unattended for hours that compounding is the entire operating budget.
This is the same argument the market has been having all year from the other direction. We covered frontier pricing moving up, not down, as the leading labs decided capability was worth a premium and let token counts grow. SpaceXAI is betting the premium is defensible only while the capability gap is visible, and two index points is not very visible.
How to read the scoreboard
Absolute index numbers drift, so treat them as a snapshot rather than a constant. Artificial Analysis rebases its index between versions and re-runs model endpoints as they are updated, which is why the figures here do not line up with the ones in our Opus 5 versus GPT-5.6 comparison from July, where Sol at max was listed two points lower. The ordering has been stable through those revisions; the decimals have not. It is also worth noting that the three scores compared here come from three different effort settings — Grok 4.6 at high, Opus 5 and Fable 5 at max — so this is a comparison of each lab's recommended top configuration, not of identical compute budgets.
The other model in the picture is Kimi K3, which Grok 4.6 has now passed. K3's weights are public, which makes it a different kind of competitor: Grok 4.6 undercuts the closed frontier on price, but it cannot undercut a model you can run yourself.
Availability
Grok 4.6 is live in the SpaceXAI API as grok-4.6, in Grok Build, and in Cursor across all plans, with double the included usage in Cursor and Grok Build for the first week. It is also on OpenRouter, Vercel and Cloudflare. SpaceXAI has kept the naming from the merger that folded xAI into SpaceX, so the model card and the API host now carry different brand names than the model line does.
If you are running long agentic jobs and watching the bill, Grok 4.6 is now the obvious thing to A/B against your current default — the turn-count numbers alone justify the test. If your workload is hard, terminal-heavy software engineering, the two evals where it trails by eight points are precisely the ones that predict your experience, and the cheaper model will not be the cheaper outcome.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.