Grok 4.7 Review: Same Price, Twice the Cost per Task

Grok 4.7 keeps Grok 4.6’s $2/$6 list price and beats it on every benchmark SpaceXAI published. The bill is another matter. Artificial Analysis measured $3.74 per Intelligence Index task at xhigh, about twice Grok 4.6’s $1.86 at high, for a score of 46 against 44. Claude Opus 5.5 at high scores 54 for $1.82. In our own build tests Grok 4.7 finished last in both rounds it entered.

By BitsMindsHands-on reviewPublished Not re-verified since

The short version

Get Grok if

  • Teams on Grok 4.6: the rates are identical and Grok 4.7 is better on every row SpaceXAI published
  • Electrical-engineering and legal agent work, the two benchmarks where it leads SpaceXAI’s table outright
  • Professional knowledge work, where it trails only Claude Fable 5.1 on SpaceXAI’s two Elo tests

Skip it if

  • Anyone buying on value: GPT-6 Sol, Muse Spark 1.3 and Claude Opus 5.5 at high all score higher for less per task
  • Front-end, animation and game work, where all four of our builds shipped a fault you could see on screen
  • Offline batch pipelines and prompts above 200K tokens, which pay the full rate or double it

What we tested

  • Four build briefs in two rounds, one attempt each, in Grok Build at xhigh (the top of its dial), against Claude Fable 5.1, GPT-6 Astra, Claude Opus 5 and GPT-5.6 Sol
  • The Artificial Analysis Intelligence Index v4.3.2 table, captured on 22 September 2026
  • SpaceXAI’s launch post and API documentation
Not re-verified since

Pros

  • Same $2/$6 list price as Grok 4.6 — with cached input at $0.50
  • Beats Grok 4.6 on every benchmark SpaceXAI published
  • Terminal-Bench 4.0 nearly doubles from 20.3% to 38.0%
  • Leads SpaceXAI’s table outright on EEBench at 64.0% and Harvey Legal Agent at 19.6%
  • Second only to Claude Fable 5.1 on both professional-work Elo tests
  • 500K-token context window with image input and structured outputs
  • Carefully built mechanisms underneath — including a real two-level ramp handoff
  • Available from launch in Cursor and Grok Build as well as the API and third-party routers and clouds

Cons

  • $3.74 per index task at xhigh — about twice Grok 4.6 for two more points
  • GPT-6 Astra and Claude Opus 5.5 at high both score higher for less per task
  • Trails Claude Fable 5.1 by nearly 20 points on Terminal-Bench 4.0
  • Finished last in both of our hands-on rounds and by wide margins
  • In all four of our briefs the visible layer was not joined to a correct mechanism
  • Slowest entrant on three of our four briefs — 80 minutes on the Hill Climb game
  • Prompts of 200K tokens or more bill every token at twice the rate
  • No Batch API and so no discount for offline work

Grok 4.7 arrived on 21 September with a pitch built on price. SpaceXAI — xAI, part of SpaceX since the April merger — kept Grok 4.6’s rates exactly: $2 per million input tokens and $6 per million output. It says the model sits on a new, larger base model, trained with a longer reinforcement-learning run weighted toward tasks that take hours. And it led the launch with a cost-per-task chart rather than a leaderboard.

That pitch can now be checked. Artificial Analysis has scored the model, and we have run it through four build briefs of our own. In July our Grok 4.5 review called SpaceXAI’s model the best value in frontier coding, on launch numbers alone. This review has independent figures and hands-on results, and they do not support the same verdict.

Same price, twice the bill

Artificial Analysis runs the same evaluation suite on every model and publishes what one task cost to complete. That figure is the list price multiplied by the tokens a model actually spends, and it is how this site judges value. It is also where Grok 4.7’s case comes apart:

Model and effortIntelligenceCost per taskList price
Claude Opus 5.5, max58$5.98$4 / $20
Claude Opus 5.5, high54$1.82$4 / $20
Claude Fable 5.1, max53$7.63$10 / $50
GPT-6 Astra, max53$3.26$10 / $50
GPT-6 Sol, max48$1.06$2 / $10
Muse Spark 1.3, max48$1.60$1.25 / $4.25
GPT-5.6 Sol, max47$1.99$4 / $20
Grok 4.7, xhigh46$3.74$2 / $6
Grok 4.6, high44$1.86$2 / $6
Cost of one Intelligence Index task, ordered by scoreUSD per v4.3.2 task, rounded · score in brackets · lower is better · Artificial AnalysisCost per task (USD)012341.8Opus 5.5 high(54)3.3Astra max (53)1.1GPT-6 Sol max(48)1.6Muse Spark 1.3(48)2.0GPT-5.6 Sol (47)3.7Grok 4.7 xhigh(46)1.9Grok 4.6 high(44)
Grok 4.7’s bar is the tallest in the chart and its score the second lowest. Every model here apart from Grok 4.6 scores higher and finishes a task for less. Grok 4.7 is measured at xhigh and Grok 4.6 at high, so part of the gap between them is the setting. Data: Artificial Analysis.

The sticker did not move and the bill did. Grok 4.7 costs $3.74 per index task against Grok 4.6’s $1.86, about twice as much, for a score of 46 against 44. The per-token rates are identical, so the difference is in how many tokens it spends to finish.

The two rows are not measured at the same setting. Artificial Analysis lists Grok 4.7 at xhigh, the top of its dial, and Grok 4.6 at high. A higher effort setting spends more tokens by design, so part of that doubling belongs to the dial rather than the model, and these figures cannot say how much. It also means the $3.74 describes Grok 4.7 at full stretch. The API defaults to high, and a call that leaves effort unset does not run at the configuration measured here.

Against the rest of the field, xhigh is hard to justify. Every model in the chart apart from Grok 4.6 scores higher than Grok 4.7 and finishes a task for less. GPT-6 Sol reaches 48 for $1.06, 28% of Grok 4.7’s bill. Claude Opus 5.5 at high scores eight points more for less than half. Even GPT-6 Astra, which lists at five times Grok’s input price, finishes a task for $3.26 and scores seven points higher.

Grok 4.7 is not the most expensive model per task. Claude Fable 5.1 costs $7.63 and Opus 5.5 at max $5.98, and both score well above it. What stands out is the price of its score. Seven models score higher on our leaderboard, and Xiaomi’s MiMo-V2.6-Pro matches its 46 for $0.13 a task. Our August coverage of Grok 4.6 quoted a score of 61 on an earlier version of the index. Versions are not comparable, and every figure here is from v4.3.2.

SpaceXAI’s benchmarks, including the ones it loses

The launch table is SpaceXAI’s own, run at effort levels it chose: Grok 4.7 at xhigh, Grok 4.6 at high, and GPT-5.6 Sol and Claude Fable 5.1 at max. Its DeepSWE figure for Grok 4.7 is a high-effort score. It was published the day before Claude Opus 5.5 launched, so Opus 5.5 is not in it.

SpaceXAI’s published benchmarksScore, % — higher is better · effort levels differ, as labelled · Grok 4.7 DeepSWE is a high-effort score · SpaceXAIGrok 4.7 (xhigh)Grok 4.6 (high)GPT-5.6 Sol (max)Claude Fable 5.1 (max)02040608046.340.441.751.8CursorBench 4.071.065.272.770.0DeepSWE v1.138.020.337.357.9Terminal-Bench4.064.053.039.456.4EEBench19.615.82.56.7Harvey LegalAgent56.748.560.562.1HealthBenchProfessional
Grok 4.7 beats Grok 4.6 on all six rows and leads two of them outright. Claude Fable 5.1 leads three and GPT-5.6 Sol one. The table predates Claude Opus 5.5, which launched a day later. Data: SpaceXAI.

Against its predecessor the gains are large on every row. Terminal-Bench 4.0, the closest of these to an agent working in a real shell, nearly doubles from 20.3% to 38.0%, and CursorBench 4.0 rises from 40.4% to 46.3%. Against the frontier the picture changes. Fable 5.1 leads CursorBench by 5.5 points and Terminal-Bench by 19.9, GPT-5.6 Sol edges DeepSWE, and on HealthBench Professional Grok 4.7’s 56.7% trails Sol’s 60.5% and Fable 5.1’s 62.1%.

Two rows go to Grok 4.7 outright. On EEBench, an electrical-engineering set, it scores 64.0% to Fable 5.1’s 56.4%. On the Harvey Legal Agent benchmark it scores 19.6% to Fable 5.1’s 6.7% and Sol’s 2.5%, close to three times the best rival on a test where every score is low. If your work sits in either field, those two rows are the reason to try it.

Professional-work Elo, points against Grok 4.7GDPval: Grok 4.7 = 1695 · AA Briefcase v1.1: Grok 4.7 = 1657 · figures published by SpaceXAIahead of Grok 4.7behind Grok 4.7-2000+200Fable 5.1 · GDPval+40Fable 5.1 · Briefcase+21Grok 4.6 · GDPval−90Grok 4.6 · Briefcase−111GPT-6 Astra · GDPval−153GPT-5.6 Sol · Briefcase−170
Only Claude Fable 5.1 finishes ahead of Grok 4.7 on either test, by 40 and 21 points. Grok 4.7 is 90 and 111 points clear of Grok 4.6. These are the figures SpaceXAI published, not independent runs.

Professional work is where it comes closest to the top. On AA Briefcase v1.1 it reaches 1657 Elo, 21 behind Fable 5.1 and 170 ahead of GPT-5.6 Sol. On GDPval it is 40 behind Fable 5.1 and 153 ahead of GPT-6 Astra. Both sets of figures come from SpaceXAI’s launch table.

What we found running it ourselves

Grok 4.7 has entered two rounds of our build-off series, both times inside Grok Build, xAI’s own agentic CLI, at xhigh, which is the top of its dial. Both rounds were scored by BitsMinds from a side-by-side page of the untouched builds, with no blind scoring and no automated grading.

The first round, published on 21 September, gave it the three briefs Claude Fable 5.1 and GPT-6 Astra had already answered: a motorway interchange and a seaside fairground as animated SVGs, and a cinematic rocket launch as one web page. One attempt each, and no looking at the result. Fable 5.1 won with 10 points, Astra 6 took 9 and Grok 4.7 finished last with 4, the widest margin the series had recorded.

Underneath, the machinery is correct and built with real care. Its cloverleaf interchange carries thirty vehicles across eight movements. Every ramp vehicle is drawn twice, once below the bridge deck and once above, and hands itself from one level to the other with matched fades at the point where its ramp crosses the deck. Its Ferris wheel’s gondolas counter-rotate so they cancel exactly. Its rocket flies the most complete flight profile the brief has produced, with booster cutoff and stage separation at T+41 and a spent booster simulated as its own falling body.

On top, the part a viewer sees was left disconnected. A painted edge line runs unbroken across the exit it should open. The wheel’s rim and bulbs stand still while the spokes turn inside them, the carousel rotates eight flat wedges of a roof that otherwise does not move, and its horses never travel. The launch camera follows the rocket upward but has no limit on how far it can fall behind sideways, so the rocket is out of frame by T+35 and the separation the sequence was built around happens off-screen. As the round’s write-up put it, in every case the mechanism is right and the layer the viewer actually sees is not joined to it.

The second round, published on 22 September, was a single brief: recreate Hill Climb Racing as one HTML file, and check the work before reporting it done. This time every entrant had a shell, Node and a browser. Claude Opus 5 won with 12 points, GPT-5.6 Sol took 10 and Grok 4.7 finished last with 6.

More of Grok’s game works than a first look suggests. The physics is sound, the wheels stay on the ground for a whole run, all seventeen pickups can be collected, and it has three real fail states. Two faults sit in the layer a player feels at once. Let go of the throttle and the buggy brakes itself, from full speed to a standstill in about half a second. And the mesas behind it are pinned to the screen rather than the ground, so their feet end up in mid-air. Its fuel tank is real, but the route hands out so much fuel that it never runs dry. Grok’s own running commentary noticed that the gauge barely moved. Its closing report said nothing in that pass was left broken.

It was slow as well. Hill Climb took it 80 minutes, the longest of the three, and it was the slowest entrant on the fairground and the launch; only its interchange, at 11 minutes 55 seconds, was the quickest of its round. Across four briefs the pattern held: Grok 4.7 builds the machine correctly and does not check what it looks like. SpaceXAI says the model verifies its own output more carefully than before. The one time it was told to check, it noticed one of its faults and shipped it anyway.

These are visual build briefs, and they test one kind of work. They say nothing about the legal and engineering tasks where SpaceXAI’s table puts it ahead. They say a good deal about front-end and game work, and about how far to trust the model’s own report that it is finished.

What changes for developers

The API id is grok-4.7 and the rates match Grok 4.6’s. SpaceXAI’s documentation has the details that decide the bill:

  • The 200K line doubles the whole request. The context window is 500,000 tokens, but a request whose prompt reaches 200,000 tokens bills every token in it at $4 input, $1.00 cached and $12 output. The surcharge is not marginal: crossing the line doubles the cost of everything in the call, not only the tokens past it.
  • Effort runs low, medium, high and xhigh, and the default is high. The independent cost figure above is for xhigh, so test at high before paying for more.
  • Caching helps; batching is not offered. Cached input costs $0.50 per million, a quarter of the normal input rate. There is no Batch API, so overnight and offline jobs pay full price.
  • A faster variant serves output at twice the speed for twice the price.
  • Input is text and images, output is text, and function calling and structured outputs are supported.

It is available in Cursor, Grok Build and the Grok API, and through third-party harnesses, routers and clouds. The pricing calculator covers the per-token side, but it multiplies list prices, and the list price is the part of Grok 4.7 that did not change.

On safety, SpaceXAI reports 62.4% on LatchBio’s biosafety benchmark and says that on HackerBench v0.3 the model lets 3.3% of risky dual-use prompts through. Selected cybersecurity partners get invite-only access to its red-team capabilities for defence research. All three are SpaceXAI’s own figures and arrangements.

Who should be running it

If you are…Do this
On Grok 4.6Switch, and start at high. The rates are the same and every published row improves; the doubling in cost per task was measured at xhigh
Doing electrical-engineering or legal agent workTest it against Claude Fable 5.1. These are the two rows where Grok 4.7 leads SpaceXAI’s table, by a wide margin on Harvey
Choosing a model on valueLook elsewhere first. GPT-6 Sol scores 48 for $1.06 a task and Claude Opus 5.5 at high scores 54 for $1.82
Building front-ends, animation or gamesLook at the output before you ship it. In all four of our briefs the fault was on the screen, not in the logic
Running long-context jobsKeep prompts under 200K tokens. From that line up, the whole request bills at double
Running offline batch pipelinesThere is no Batch API. Claude Opus 5.5’s batch tier halves its rates to $2/$10

Verdict

Grok 4.7 earns a 4.0 out of 5, below the 4.4 we gave Grok 4.5 in July on launch numbers alone, and the lowest rating on this site. It is a better model than Grok 4.6 on every row SpaceXAI published. It leads SpaceXAI’s table outright on electrical-engineering and legal agent work, and it is second only to Fable 5.1 on both professional-work Elo tests. What it builds underneath is correct and made with care.

The score falls on the two things that were meant to be its strengths. Value was the case for Grok, and at the setting Artificial Analysis measured, a task costs about twice what Grok 4.6 spent, more than GPT-6 Astra and more than twice Claude Opus 5.5 at high, for a score of 46. In our own tests it finished last twice, with the same fault both times.

It is a capable model sold on the wrong number. Run it at high, where the API default already sits, and measure the bill on your own jobs before assuming the sticker is what you will pay. If your work is in its two strong fields, test it against Fable 5.1. For most other work, the leaderboard lists several models that score higher and cost less per task.

About the score. 4.0/ 5 is BitsMinds' editorial verdict from our own testing and research — not an average of user ratings, which we do not collect. Prices and plan tiers are as published by the vendor on the fact-check date shown above. How we rate →

Related Reviews

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.