Grok 4.7 Review: Same Price, Twice the Cost per Task
Grok 4.7 keeps Grok 4.6’s $2/$6 list price and beats it on every benchmark SpaceXAI published. The bill is another matter. Artificial Analysis measured $3.74 per Intelligence Index task at xhigh, about twice Grok 4.6’s $1.86 at high, for a score of 46 against 44. Claude Opus 5.5 at high scores 54 for $1.82. In our own build tests Grok 4.7 finished last in both rounds it entered.
The short version
Get Grok if
- Teams on Grok 4.6: the rates are identical and Grok 4.7 is better on every row SpaceXAI published
- Electrical-engineering and legal agent work, the two benchmarks where it leads SpaceXAI’s table outright
- Professional knowledge work, where it trails only Claude Fable 5.1 on SpaceXAI’s two Elo tests
Skip it if
- Anyone buying on value: GPT-6 Sol, Muse Spark 1.3 and Claude Opus 5.5 at high all score higher for less per task
- Front-end, animation and game work, where all four of our builds shipped a fault you could see on screen
- Offline batch pipelines and prompts above 200K tokens, which pay the full rate or double it
What we tested
- Four build briefs in two rounds, one attempt each, in Grok Build at xhigh (the top of its dial), against Claude Fable 5.1, GPT-6 Astra, Claude Opus 5 and GPT-5.6 Sol
- The Artificial Analysis Intelligence Index v4.3.2 table, captured on 22 September 2026
- SpaceXAI’s launch post and API documentation
Pros
- Same $2/$6 list price as Grok 4.6 — with cached input at $0.50
- Beats Grok 4.6 on every benchmark SpaceXAI published
- Terminal-Bench 4.0 nearly doubles from 20.3% to 38.0%
- Leads SpaceXAI’s table outright on EEBench at 64.0% and Harvey Legal Agent at 19.6%
- Second only to Claude Fable 5.1 on both professional-work Elo tests
- 500K-token context window with image input and structured outputs
- Carefully built mechanisms underneath — including a real two-level ramp handoff
- Available from launch in Cursor and Grok Build as well as the API and third-party routers and clouds
Cons
- $3.74 per index task at xhigh — about twice Grok 4.6 for two more points
- GPT-6 Astra and Claude Opus 5.5 at high both score higher for less per task
- Trails Claude Fable 5.1 by nearly 20 points on Terminal-Bench 4.0
- Finished last in both of our hands-on rounds and by wide margins
- In all four of our briefs the visible layer was not joined to a correct mechanism
- Slowest entrant on three of our four briefs — 80 minutes on the Hill Climb game
- Prompts of 200K tokens or more bill every token at twice the rate
- No Batch API and so no discount for offline work
Grok 4.7 arrived on 21 September with a pitch built on price. SpaceXAI — xAI, part of SpaceX since the April merger — kept Grok 4.6’s rates exactly: $2 per million input tokens and $6 per million output. It says the model sits on a new, larger base model, trained with a longer reinforcement-learning run weighted toward tasks that take hours. And it led the launch with a cost-per-task chart rather than a leaderboard.
That pitch can now be checked. Artificial Analysis has scored the model, and we have run it through four build briefs of our own. In July our Grok 4.5 review called SpaceXAI’s model the best value in frontier coding, on launch numbers alone. This review has independent figures and hands-on results, and they do not support the same verdict.
Same price, twice the bill
Artificial Analysis runs the same evaluation suite on every model and publishes what one task cost to complete. That figure is the list price multiplied by the tokens a model actually spends, and it is how this site judges value. It is also where Grok 4.7’s case comes apart:
| Model and effort | Intelligence | Cost per task | List price |
|---|---|---|---|
| Claude Opus 5.5, max | 58 | $5.98 | $4 / $20 |
| Claude Opus 5.5, high | 54 | $1.82 | $4 / $20 |
| Claude Fable 5.1, max | 53 | $7.63 | $10 / $50 |
| GPT-6 Astra, max | 53 | $3.26 | $10 / $50 |
| GPT-6 Sol, max | 48 | $1.06 | $2 / $10 |
| Muse Spark 1.3, max | 48 | $1.60 | $1.25 / $4.25 |
| GPT-5.6 Sol, max | 47 | $1.99 | $4 / $20 |
| Grok 4.7, xhigh | 46 | $3.74 | $2 / $6 |
| Grok 4.6, high | 44 | $1.86 | $2 / $6 |
The sticker did not move and the bill did. Grok 4.7 costs $3.74 per index task against Grok 4.6’s $1.86, about twice as much, for a score of 46 against 44. The per-token rates are identical, so the difference is in how many tokens it spends to finish.
The two rows are not measured at the same setting. Artificial Analysis lists Grok 4.7 at xhigh, the top of its dial, and Grok 4.6 at high. A higher effort setting spends more tokens by design, so part of that doubling belongs to the dial rather than the model, and these figures cannot say how much. It also means the $3.74 describes Grok 4.7 at full stretch. The API defaults to high, and a call that leaves effort unset does not run at the configuration measured here.
Against the rest of the field, xhigh is hard to justify. Every model in the chart apart from Grok 4.6 scores higher than Grok 4.7 and finishes a task for less. GPT-6 Sol reaches 48 for $1.06, 28% of Grok 4.7’s bill. Claude Opus 5.5 at high scores eight points more for less than half. Even GPT-6 Astra, which lists at five times Grok’s input price, finishes a task for $3.26 and scores seven points higher.
Grok 4.7 is not the most expensive model per task. Claude Fable 5.1 costs $7.63 and Opus 5.5 at max $5.98, and both score well above it. What stands out is the price of its score. Seven models score higher on our leaderboard, and Xiaomi’s MiMo-V2.6-Pro matches its 46 for $0.13 a task. Our August coverage of Grok 4.6 quoted a score of 61 on an earlier version of the index. Versions are not comparable, and every figure here is from v4.3.2.
SpaceXAI’s benchmarks, including the ones it loses
The launch table is SpaceXAI’s own, run at effort levels it chose: Grok 4.7 at xhigh, Grok 4.6 at high, and GPT-5.6 Sol and Claude Fable 5.1 at max. Its DeepSWE figure for Grok 4.7 is a high-effort score. It was published the day before Claude Opus 5.5 launched, so Opus 5.5 is not in it.
Against its predecessor the gains are large on every row. Terminal-Bench 4.0, the closest of these to an agent working in a real shell, nearly doubles from 20.3% to 38.0%, and CursorBench 4.0 rises from 40.4% to 46.3%. Against the frontier the picture changes. Fable 5.1 leads CursorBench by 5.5 points and Terminal-Bench by 19.9, GPT-5.6 Sol edges DeepSWE, and on HealthBench Professional Grok 4.7’s 56.7% trails Sol’s 60.5% and Fable 5.1’s 62.1%.
Two rows go to Grok 4.7 outright. On EEBench, an electrical-engineering set, it scores 64.0% to Fable 5.1’s 56.4%. On the Harvey Legal Agent benchmark it scores 19.6% to Fable 5.1’s 6.7% and Sol’s 2.5%, close to three times the best rival on a test where every score is low. If your work sits in either field, those two rows are the reason to try it.
Professional work is where it comes closest to the top. On AA Briefcase v1.1 it reaches 1657 Elo, 21 behind Fable 5.1 and 170 ahead of GPT-5.6 Sol. On GDPval it is 40 behind Fable 5.1 and 153 ahead of GPT-6 Astra. Both sets of figures come from SpaceXAI’s launch table.
What we found running it ourselves
Grok 4.7 has entered two rounds of our build-off series, both times inside Grok Build, xAI’s own agentic CLI, at xhigh, which is the top of its dial. Both rounds were scored by BitsMinds from a side-by-side page of the untouched builds, with no blind scoring and no automated grading.
The first round, published on 21 September, gave it the three briefs Claude Fable 5.1 and GPT-6 Astra had already answered: a motorway interchange and a seaside fairground as animated SVGs, and a cinematic rocket launch as one web page. One attempt each, and no looking at the result. Fable 5.1 won with 10 points, Astra 6 took 9 and Grok 4.7 finished last with 4, the widest margin the series had recorded.
Underneath, the machinery is correct and built with real care. Its cloverleaf interchange carries thirty vehicles across eight movements. Every ramp vehicle is drawn twice, once below the bridge deck and once above, and hands itself from one level to the other with matched fades at the point where its ramp crosses the deck. Its Ferris wheel’s gondolas counter-rotate so they cancel exactly. Its rocket flies the most complete flight profile the brief has produced, with booster cutoff and stage separation at T+41 and a spent booster simulated as its own falling body.
On top, the part a viewer sees was left disconnected. A painted edge line runs unbroken across the exit it should open. The wheel’s rim and bulbs stand still while the spokes turn inside them, the carousel rotates eight flat wedges of a roof that otherwise does not move, and its horses never travel. The launch camera follows the rocket upward but has no limit on how far it can fall behind sideways, so the rocket is out of frame by T+35 and the separation the sequence was built around happens off-screen. As the round’s write-up put it, in every case the mechanism is right and the layer the viewer actually sees is not joined to it.
The second round, published on 22 September, was a single brief: recreate Hill Climb Racing as one HTML file, and check the work before reporting it done. This time every entrant had a shell, Node and a browser. Claude Opus 5 won with 12 points, GPT-5.6 Sol took 10 and Grok 4.7 finished last with 6.
More of Grok’s game works than a first look suggests. The physics is sound, the wheels stay on the ground for a whole run, all seventeen pickups can be collected, and it has three real fail states. Two faults sit in the layer a player feels at once. Let go of the throttle and the buggy brakes itself, from full speed to a standstill in about half a second. And the mesas behind it are pinned to the screen rather than the ground, so their feet end up in mid-air. Its fuel tank is real, but the route hands out so much fuel that it never runs dry. Grok’s own running commentary noticed that the gauge barely moved. Its closing report said nothing in that pass was left broken.
It was slow as well. Hill Climb took it 80 minutes, the longest of the three, and it was the slowest entrant on the fairground and the launch; only its interchange, at 11 minutes 55 seconds, was the quickest of its round. Across four briefs the pattern held: Grok 4.7 builds the machine correctly and does not check what it looks like. SpaceXAI says the model verifies its own output more carefully than before. The one time it was told to check, it noticed one of its faults and shipped it anyway.
These are visual build briefs, and they test one kind of work. They say nothing about the legal and engineering tasks where SpaceXAI’s table puts it ahead. They say a good deal about front-end and game work, and about how far to trust the model’s own report that it is finished.
What changes for developers
The API id is grok-4.7 and the rates match Grok 4.6’s. SpaceXAI’s documentation has the details that decide the bill:
- The 200K line doubles the whole request. The context window is 500,000 tokens, but a request whose prompt reaches 200,000 tokens bills every token in it at $4 input, $1.00 cached and $12 output. The surcharge is not marginal: crossing the line doubles the cost of everything in the call, not only the tokens past it.
- Effort runs low, medium, high and xhigh, and the default is high. The independent cost figure above is for xhigh, so test at high before paying for more.
- Caching helps; batching is not offered. Cached input costs $0.50 per million, a quarter of the normal input rate. There is no Batch API, so overnight and offline jobs pay full price.
- A faster variant serves output at twice the speed for twice the price.
- Input is text and images, output is text, and function calling and structured outputs are supported.
It is available in Cursor, Grok Build and the Grok API, and through third-party harnesses, routers and clouds. The pricing calculator covers the per-token side, but it multiplies list prices, and the list price is the part of Grok 4.7 that did not change.
On safety, SpaceXAI reports 62.4% on LatchBio’s biosafety benchmark and says that on HackerBench v0.3 the model lets 3.3% of risky dual-use prompts through. Selected cybersecurity partners get invite-only access to its red-team capabilities for defence research. All three are SpaceXAI’s own figures and arrangements.
Who should be running it
| If you are… | Do this |
|---|---|
| On Grok 4.6 | Switch, and start at high. The rates are the same and every published row improves; the doubling in cost per task was measured at xhigh |
| Doing electrical-engineering or legal agent work | Test it against Claude Fable 5.1. These are the two rows where Grok 4.7 leads SpaceXAI’s table, by a wide margin on Harvey |
| Choosing a model on value | Look elsewhere first. GPT-6 Sol scores 48 for $1.06 a task and Claude Opus 5.5 at high scores 54 for $1.82 |
| Building front-ends, animation or games | Look at the output before you ship it. In all four of our briefs the fault was on the screen, not in the logic |
| Running long-context jobs | Keep prompts under 200K tokens. From that line up, the whole request bills at double |
| Running offline batch pipelines | There is no Batch API. Claude Opus 5.5’s batch tier halves its rates to $2/$10 |
Verdict
Grok 4.7 earns a 4.0 out of 5, below the 4.4 we gave Grok 4.5 in July on launch numbers alone, and the lowest rating on this site. It is a better model than Grok 4.6 on every row SpaceXAI published. It leads SpaceXAI’s table outright on electrical-engineering and legal agent work, and it is second only to Fable 5.1 on both professional-work Elo tests. What it builds underneath is correct and made with care.
The score falls on the two things that were meant to be its strengths. Value was the case for Grok, and at the setting Artificial Analysis measured, a task costs about twice what Grok 4.6 spent, more than GPT-6 Astra and more than twice Claude Opus 5.5 at high, for a score of 46. In our own tests it finished last twice, with the same fault both times.
It is a capable model sold on the wrong number. Run it at high, where the API default already sits, and measure the bill on your own jobs before assuming the sticker is what you will pay. If your work is in its two strong fields, test it against Fable 5.1. For most other work, the leaderboard lists several models that score higher and cost less per task.
About the score. 4.0/ 5 is BitsMinds' editorial verdict from our own testing and research — not an average of user ratings, which we do not collect. Prices and plan tiers are as published by the vendor on the fact-check date shown above. How we rate →
Related Reviews
Claude Opus 5.5 takes sole first place on the Artificial Analysis Intelligence Index, scoring 58 to the 53 shared by Claude Fable 5.1 and GPT-6 Astra. It lists at $4/$20, a fifth under Opus 5. The better story sits one effort level down: at high it scores 54, above every rival, for $1.82 a task. It also won our three-round build-off, and it was the slowest entrant in every round.
Read review →Anthropic turned the lineup over again: Fable 5.1 replaced Fable 5 at the top on September 1, Opus 5 replaced Opus 4.8 at the same price, and Sonnet 5’s intro pricing has ended. Where Claude leads, where it still struggles, and whether Pro and Max are worth it in September 2026.
Read review →ChatGPT Plus is still the best $20 in AI and Pro is still a $200 question — now with GPT-6 Astra, Sol and Luna rolling out to both tiers. What changed since our GPT-5.5 hands-on, where Claude still wins, and who actually needs Pro.
Read review →Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.