Grok 4.7 Chases Fable 5.1 at a Fraction of the Price
SpaceXAI released Grok 4.7 on 21 September at exactly the price Grok 4.6 carried — $2 per million input tokens and $6 per million output. It beats its predecessor on every benchmark the company published, trails Claude Fable 5.1 on most of them, and costs a fifth as much to feed and an eighth as much to read.
SpaceXAI released Grok 4.7 on 21 September, calling it its most capable model for coding and knowledge work. The company is not asking anyone to pay more for it: the model ships at $2 per million input tokens and $6 per million output, the same rates Grok 4.6 has carried since August, with a faster variant served at twice the output speed for twice the price. It is available today in Cursor and Grok Build, and through the Grok API, third-party coding harnesses, model routers and cloud platforms.
The changes underneath are more than a tune-up. SpaceXAI says Grok 4.7 uses a new, larger base model than 4.6, trained with a longer reinforcement-learning run on a harder mix of tasks weighted toward problems that take many hours to finish. The stated result is a model that works longer on difficult tasks, verifies its own output more carefully and handles long context better. It was also trained to natively understand the Grok Bot harness, which is where the company's claim of improved conversational and general knowledge work comes from.
On the numbers, the honest summary is that Grok 4.7 beats the model it replaces everywhere and beats the frontier almost nowhere. It improves on Grok 4.6 across every benchmark SpaceXAI published — CursorBench 4.0 rises from 40.4 to 46.3, Terminal-Bench 4.0 from 20.3 to 38.0 — but Claude Fable 5.1 still leads CursorBench (51.8) and Terminal-Bench (57.9) by wide margins. Two caveats belong on every figure here: they are vendor-reported, and the effort levels are not matched, with Grok 4.7 run at xHigh against rivals at Max. The DeepSWE v1.1 score of 71.0 is separately flagged by the company as a high-effort result.
There are two benchmarks where Grok 4.7 leads outright, and one of them is not close. On EEBench, an electrical engineering set, it scores 64.0 against 56.4 for Fable 5.1 and 39.4 for GPT-5.6 Sol. On the Harvey Legal Agent Benchmark it scores 19.6 against Fable 5.1's 6.7 and Sol's 2.5 — close to three times the best rival score, on a benchmark where every number is low enough to suggest the task is far from solved. On clinical reasoning it loses: HealthBench Professional puts it at 56.7 behind Sol's 60.5 and Fable 5.1's 62.1.
| Benchmark | Grok 4.7 (xhigh) | Grok 4.6 (high) | GPT-5.6 Sol (max) | Fable 5.1 (max) |
|---|---|---|---|---|
| Input price, $/M tokens | $2 | $2 | $4 | $10 |
| Output price, $/M tokens | $6 | $6 | $20 | $50 |
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
| Harvey Legal Agent | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
| AA Briefcase v1.1 (Elo) | 1,657 | 1,546 | 1,487 | 1,678 |
The framing SpaceXAI chose for the launch is the one that matters, and it is not a leaderboard position. The company led with a cost-per-task chart on CursorBench 4.0 rather than a score chart, plotting average dollars spent to complete a task against accuracy — and on that axis Grok 4.7 sits at the frontier. That is the right measure for a reasoning model, and it is the measure on which the gap closes: Fable 5.1 scores 5.5 points higher on CursorBench while listing at five times the input price and more than eight times the output price. On GDPval, a professional-work Elo benchmark, Fable 5.1 leads at 1,735 to Grok 4.7's 1,695, with Grok 4.6 at 1,605 and GPT-6 Astra at 1,542.
Safety is the section SpaceXAI gave the most new ground to. The company says Grok 4.7 was built on an entirely new safeguard stack and is the strongest model it has tested on refusals and jailbreak resistance. It reports topping LatchBio's biosafety benchmark at 62.4%, and on HackerBench v0.3 — its own benchmark for risky and malicious cyber tasks — says the model allows only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. SpaceXAI has also begun giving selected cybersecurity partners invite-only access to Grok 4.7's red-team capabilities for defence research, an arrangement that puts it alongside the defenders-only cyber variants Google and Anthropic have shipped this month.
What SpaceXAI has not published is an independent number. Every figure above is the company's own, run at an effort level it chose, against rivals it configured. Artificial Analysis has not yet scored Grok 4.7, and until it does — particularly its cost-to-run-the-index figure, which counts output tokens rather than list prices — the price-performance claim rests on a chart drawn by the party making it.
More on Grok
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.