Part of our AI Comparisons
Models·15 min read
By BitsMinds Hands-on test

Grok 4.7 vs Fable 5.1 vs Astra 6: Can It Compete?

SpaceXAI’s Grok 4.7 took the three briefs this series has handed every frontier model — a motorway interchange, a seaside fairground and a cinematic rocket launch — against Claude Fable 5.1 and GPT-6 Astra. One attempt each, maximum reasoning effort, no looking at what they built. Every mechanism Grok was asked for is present and computed correctly. In all three, the thing a viewer actually watches was never connected to it.

Grok 4.7 Astra 6 Fable 5.1 VS VS BITSMINDS.COM
Share:
BitsMinds LabHow this test was run3 models · 3 briefs · one attempt per model per brief · Maximum reasoning effort for every model. Grok 4.7 ran at Grok Build's xhigh setting, the same level Astra was given in CodexShow details
Models
Claude Fable 5.1 (Claude Code) · GPT-6 Astra (OpenAI Codex) · Grok 4.7 (Grok Build)
Attempts
One attempt per model per brief. No retries, no follow-up notes, no cleanup.
Reasoning effort
Maximum reasoning effort for every model. Grok 4.7 ran at Grok Build's xhigh setting, the same level Astra was given in Codex.
Tools allowed
Each model ran inside its own maker's agentic coding environment — Fable 5.1 in Claude Code, Astra in Codex, Grok 4.7 in Grok Build — with the same ban on opening a browser to check its own work.
Timing
Wall-clock time from prompt to finished file, measured for every round.
Scoring
4 points for a round win, scored on the finished artefact. Fable 5.1 and Astra 6 carry their scores from the September head-to-head on these same three briefs; Grok 4.7 is scored into that set. Final: Fable 5.1 10, Astra 6 9, Grok 4.7 4.
Judging
Scored by BitsMinds from a local side-by-side comparison page of the untouched outputs. No blind scoring, no automated grading.
Published
September 21, 2026

Result

ModelProviderScore
Claude Fable 5.1Anthropic10 pts
GPT-6 AstraOpenAI9 pts
Grok 4.7SpaceXAI4 pts

10–9–4. Every mechanism the briefs asked Grok 4.7 for is present and computed correctly; in all three rounds the layer the viewer actually sees was left disconnected from it.

The briefs — as described in the article; the exact prompt files were not published

  1. Round 1: Motorway interchange (SVG)

    A top-down view of a busy interchange as one self-contained SVG. Two roads crossing on different levels, connecting ramps, at least twelve vehicles following the curve of the road they are on, and correct layering so traffic on the lower road is hidden by the overpass.

    Result: Astra 4, Fable 5.1 2, Grok 4.7 2: Grok built the best-engineered ramps in the series and ran its edge line straight through the exit.

  2. Round 2: Seaside fairground (SVG)

    A seaside fairground at dusk, one SVG. A Ferris wheel with at least ten gondolas turning continuously, a spinning carousel with at least six horses rising and falling out of phase, twinkling bulbs, a reflection in the water. Each gondola must stay level with the horizon throughout the rotation.

    Result: Fable 5.1 4, Astra 3, Grok 4.7 1: Grok's wheel rim never turns and its carousel horses never travel.

  3. Round 3: Rocket launch (web page)

    A self-contained web page that animates a rocket launching — countdown, ignition, liftoff and ascent — with a mission-control HUD. No library named, no 3D requested, one aesthetic instruction: make it look good.

    Result: Fable 5.1 4, Astra 2, Grok 4.7 1: the most complete flight profile in the series, played to an empty frame.

Read with care

  • Single run per model — a re-run could land differently.
  • Scored on the finished artefact only. Wall-clock time is reported but carries no points.
  • Three different harnesses (Claude Code, Codex, Grok Build) mean the environment, not just the model, was under test.

Part of BitsMinds Lab, our series of original hands-on tests.

Since June this series has handed the same build briefs to frontier models and published whatever came back, untouched. Two labs have taken part so far. This round adds a third: SpaceXAI, whose Grok 4.7 ran the three briefs from the September cross-lab round against the two models that have already answered them — Anthropic’s Claude Fable 5.1 and OpenAI’s GPT-6 Astra.

Three briefs: a top-down motorway interchange and a seaside fairground, both as single animated SVG files, and a cinematic rocket launch as one web page. One attempt each. Maximum reasoning effort. No opening a browser to check the result. Grok 4.7 ran inside Grok Build, xAI’s own agentic CLI, at its xhigh setting — the same effort level Astra was given in Codex. All nine builds run live below.

Grok 4.7 came last, and not narrowly: 4 points against Fable 5.1’s 10 and Astra 6’s 9, the widest margin this series has recorded. But the way it lost is more interesting than the scoreline, because it lost the same way three times.

Every mechanism the briefs asked for is in its files, and most are built with real care. Its ramps genuinely climb between two road levels. Its Ferris wheel’s counter-rotation cancels to the frame. Its rocket separates its stages and models the spent booster as a falling body with its own physics. And in all three builds, something the viewer actually looks at was left disconnected from the machinery underneath it: a painted line drawn over the exit it should open, a ride whose shell never turns with its core, a camera that loses the rocket before the moment it exists to show.

Round 1: The interchange

The brief: a top-down view of a busy interchange as one self-contained SVG. Two roads crossing on different levels, connecting ramps, at least twelve vehicles following the curve of the road they are on, and correct layering so traffic on the lower road is hidden by the overpass. It carries two traps that go unstated — whether the ramps genuinely connect the two roads, and whether traffic flows the right way after merging. Both have caught models before.

Animated motorway interchange built in one attempt by Grok 4.7
Grok 4.7, 11 minutes 55 seconds. A full cloverleaf: thirty vehicles across eight movements, four of them 270-degree loop ramps.
Animated motorway interchange built in one attempt by GPT-6 Astra
Astra 6, 16 minutes 15 seconds. Forty-eight vehicles, the most of any entry in the series, with headlight cones and direction arrows painted on the asphalt.
Animated motorway interchange built in one attempt by Claude Fable 5.1
Fable 5.1, 25 minutes 55 seconds. Still the only interchange anyone has built for this series with no visible fault.

Grok 4.7 cleared both hidden traps, and cleared them well. It built a complete cloverleaf: four through movements and four 270-degree loop ramps, thirty vehicles in all. Every ramp leaves from the outer right-hand lane of the road it exits and arrives in the outer right-hand lane of the road it joins, consistently right-hand drive in all eight movements. Every one of the thirty vehicle paths begins and ends off-screen, so the loop restarts where nobody can see it.

The layering is not a trick either. The eight vehicles on the lower road are painted before the bridge deck and the six on the upper road after it, so traffic really does vanish underneath. The deck carries expansion joints at exactly the two points where it spans the road below, which is where a real bridge puts them, and it casts a shadow onto the carriageway on both approaches.

Best of all is the part nobody asked for. A ramp climbing from the lower road to the upper one has to change levels, and Grok 4.7 solved it by drawing every ramp vehicle twice — once below the deck, once above — with complementary opacity animations on identical key times. The lower copy fades out over 72 milliseconds while the upper copy fades in at the same position, and the car hands itself from one level to the other. The two climbing ramps switch at 60% of their path, the two descending ramps at 47%, each matched to where that ramp actually crosses the deck. It is the most carefully wired mechanism anyone has submitted to this brief.

And then the road is painted over it. The carriageway’s solid white edge line is a single line element running the full width of the scene with no break in it, and it is drawn after the ramps, on top of them. The ramp crosses it 47.6% of the way along its path and the line does not part: there is no gore, no taper, no opening. The exit is a lane that simply passes under an unbroken painted line.

Astra 6 built the richest interchange anyone has drawn for this series, in 16 minutes 15 seconds. Forty-eight vehicles, more than any other entry, every one throwing a headlight cone. Direction arrows are painted onto the road surface, which has the useful side effect of making the traffic-direction trap checkable by eye. There is street lighting down the median, road signage, an articulated lorry and a bus in the traffic. Its bridge is a real structure: a deck with a shadow beneath it, parapets, bearings at all four corners, and lane markings that stop cleanly at the deck edge. Its occlusion works mid-crossing — an articulated truck enters from the west with its cab already hidden behind the parapet while its trailer is still in the open — and every one of its vehicle paths begins off-screen, so ramp traffic arrives from the mainline rather than appearing out of nothing.

And its white shoulder line runs straight through the exit, so its ramp is never visually separated from the road it leaves either. Claude Opus 5 made the same mistake in the four-model Claude run. Grok 4.7 makes it here. Three models, three labs, one detail.

Fable 5.1 is the plainest of the three and the only one with nothing wrong with it. It took 25 minutes 55 seconds and has neither Astra’s density nor Grok’s ramp engineering: no headlight cones, no painted arrows, no street furniture. What it has is an interchange whose exits separate from the carriageway properly, whose traffic stays on the correct side after every merge, and which hides exactly the right vehicles under the bridge. It is the only build in the series that has never given a reader anything to catch.

Round 1 to Astra 6, 4 points. It is faster and far denser than Fable 5.1, and its single flaw was too small to hand the round to a quieter entry. Fable 5.1 takes 2 for a build with nothing wrong with it. Grok 4.7 takes 2 as well — the fastest interchange of the three at 11 minutes 55 seconds and the best-engineered ramps in the series, landing level with the fault-free build because the one thing it got wrong is the thing that keeps being got wrong.

Round 2: The fairground

A seaside fairground at dusk, one SVG. A Ferris wheel with at least ten gondolas turning continuously, each staying level with the horizon throughout the rotation. A carousel beside it: a spinning canopy with at least six horses rising and falling out of phase. Twinkling bulbs, lights chasing the wheel’s rim, a dusk sky and a reflection in the water.

Animated seaside fairground at dusk built in one attempt by Grok 4.7
Grok 4.7, 37 minutes 46 seconds. A handsome scene — and three rides with the same thing wrong with them.
Animated seaside fairground at dusk built in one attempt by GPT-6 Astra
Astra 6, 26 minutes 33 seconds. The best-drawn carousel horses in the series, on a ride whose canopy never turns.
Animated seaside fairground at dusk built in one attempt by Claude Fable 5.1
Fable 5.1, 35 minutes 16 seconds. Its carousel turns.

Grok 4.7’s scene is genuinely attractive: a graded dusk sky, a low sun on the water, string lights between the promenade poles, silhouetted figures on the boardwalk, a booth, balloons, birds, and the whole fairground mirrored in rippling water along the bottom. There are 111 separate opacity animations across 27 different durations and 23 staggered start offsets, which is the twinkle and the chase, and they work.

The stated trap is the gondolas, and all three models cleared it. Grok 4.7 cleared it exactly: twelve gondolas, the wheel turning a full circle every 22 seconds, each gondola counter-rotating a full circle over the same 22 seconds so they cancel precisely, with a gentle sway layered on top at four different periods so they never move as one.

Then look at what its wheel is made of. The spokes and the gondolas sit inside the rotating group. The rim does not. The outer hoop, its two gold highlight rings, the inner ring and the ring of bulbs are all drawn outside the rotation and never move. The spokes turn underneath a hoop that stands still, which on a real Ferris wheel is not a thing that can happen: the rim is welded to the spokes.

The carousel has the same fault twice over. Its canopy rotation is real — a full 360 degrees every 12 seconds — but it is applied to eight flat pie wedges and nothing else. The cone of the roof, its scalloped edge, the rim and the platform are all static. The stripes revolve inside a roof that does not.

And the horses do not travel at all. There are six, they rise and fall 11 units on a 1.8-second cycle, and their start offsets are staggered −0.3, −0.6, −0.9, −1.2 and −1.5 seconds, which is the “out of phase” requirement done properly. But their horizontal positions are fixed constants. Sampled across a full canopy revolution — at 0, 1.5, 3, 6, 9 and 11.9 seconds — all six sit at the same six x-coordinates to the decimal, every time. They are arranged in a ring in perspective, scaled from 0.66 at the back to 1.1 at the front, which makes a circle that never comes round all the more visible.

Astra 6’s fairground is the best-drawn of the three, in 26 minutes 33 seconds: a pastel dusk with a lighthouse and a sailboat, and carousel horses with jewelled saddle blankets, bridles and flowing manes, the finest drawing anyone has submitted to this brief. Its Ferris wheel clears the stated trap the same way Grok’s does, turning a full circle every 48 seconds while each gondola counter-rotates 360 degrees over the same 48 seconds, cancelling exactly.

Its carousel does not turn at all. Pulling apart all 37 of its animations accounts for every one: thirteen full rotations, each belonging to the Ferris wheel; eleven vertical translations, which are the horses on their poles; twelve gondola sways; and one animation left over — a three-degree rock back and forth about the carousel’s centre, over eight seconds. That is the entire ride. The canopy rocks, the horses bob on the spot, nothing revolves, and the horses never face a direction of travel because there is no travel.

Fable 5.1’s carousel turns, and it is the only one in the round that does. It took 35 minutes 16 seconds and is warmer and busier than Astra’s, if less finely drawn. Its ride is built as one rotating assembly with the whole structure inside it — canopy, rim and horses together — so the horses travel with the roof and rise and fall as they go, which is the entire difference between a carousel and three separate animations pointed at the same spot.

This is the ride that keeps breaking, and it never breaks the same way twice: Astra’s canopy does not rotate, GPT-5.6 Sol’s rotated and left its horses bobbing on the spot, and Grok 4.7 manages both at once by rotating eight flat wedges of a roof that otherwise stands still. The trap the brief names gets solved by everybody. The trap of equal difficulty it does not name has still only ever been built correctly by a Claude model.

Round 2 to Fable 5.1, 4 points — warmer, busier, and the only carousel in the round that works as a carousel. Astra 6 takes 3 for the drawing. Grok 4.7 takes 1.

Round 3: The rocket launch

The only brief in the series that does not say how to build. It names no library, does not ask for 3D, and offers one aesthetic instruction: make it look good. A countdown, ignition, liftoff, a sustained ascent, and a mission-control HUD reporting the state of the flight.

There is a joke available here and the result declines to make it funnier: xAI has been part of SpaceX since the April merger. This is the rocket company’s model, on the rocket brief, and it is the round it lost worst.

LYRA-1 — Grok 4.7, 50 minutes 48 seconds.

Astra 6, 28 minutes.

Fable 5.1, 49 minutes 7 seconds. All three start on their own countdown.

Grok 4.7 built the most complete flight profile this brief has produced. Countdown, ignition, liftoff, a pitch program, Mach 1, booster engine cutoff and stage separation at T+41, second-stage ignition, and the Karman line at T+1:34. The cutoff is real: the throttle drops to zero and the spent booster becomes a separately simulated body that tumbles away on its own. Beside it runs a mission-control HUD reporting altitude, velocity, Mach, acceleration, dynamic pressure, pitch, propellant and engine count, with a flight-director event log.

None of which you will see. The camera follows the rocket vertically with a hard limit on how far it may fall behind, and horizontally with no limit at all. As the vehicle pitches over and builds downrange speed, that horizontal lag grows without bound: the rocket holds dead centre until about T+26, is three-quarters of the way to the right-hand edge by T+30, and is gone by T+35. So the separation at T+41 — the flash, the tumbling booster, the second-stage ignition, the whole beat the sequence is built around — happens outside the frame, and the upper stage is never seen at all. The only thing that crosses the picture afterwards is the discarded booster falling through it.

The flight also never ends. The throttle stays pinned at 100%, the propellant model decays toward 16% without ever reaching it, and the page header announces an uncrewed orbital insertion that the mission never performs. Smaller things pile up around that: clouds drawn as single flat ellipses, grid fins fixed at a permanent 24-degree cant, nothing casting a shadow anywhere, a sound design with no control to turn it on or off, and no way to watch the launch again short of reloading the page.

Astra 6 built the best-looking product of the three by some distance, and the shortest flight. A three-core vehicle rendered with real care, enormous editorial typography, the curvature of the Earth below, a flight timeline, a zoom control and a playback-speed multiplier nobody asked for. Its ascent physics are modelled rather than decorative: 38.5 km at 1,003 metres per second by T+1:39 tracks a real ascent closely, and its g-load climbing from 2.13 to 3.29 as propellant burns away is correct behaviour rather than a number chosen to look busy. And its flight sequence reads countdown, ignition, liftoff, max-Q, ascent — and stops. No cutoff, no separation, no second stage, no orbit. It is a first act with no second one.

Fable 5.1 flies the mission. The least striking image of the three and the most complete flight: it carries the staging Astra lacks, and it computes the speed of sound from altitude rather than dividing by a fixed 340 m/s, reporting Mach 0.73 at 4.54 km where the standard atmosphere gives 0.729. Its own fault is that its burn never cuts off either, and the camera judders late in the ascent. But it keeps the vehicle in shot the whole way.

Round 3 to Fable 5.1, 4 points. Astra 6 takes 2, Grok 4.7 takes 1 — the most complete flight profile in the round, played to an empty frame.

Nobody reaches for the library

One thing this brief keeps proving. It explicitly permits any library from a public CDN, and in June both entrants pulled Three.js and WebGL without being asked. Grok 4.7 did not. It hand-built the whole thing on two 2D canvases, 73 KB and 2,218 lines, with no external script tag anywhere.

That makes five models from three labs, all given written permission to use a 3D engine, all declining it. Whatever changed in what these models consider the obvious tool, it changed at OpenAI, Anthropic and xAI alike.

The clock

BuildGrok 4.7Astra 6Fable 5.1
Interchange11m 55s16m 15s25m 55s
Fairground37m 46s26m 33s35m 16s
Rocket launch50m 48s28m 00s49m 07s

Grok 4.7 produced the fastest interchange in this three-way and then the slowest fairground and the slowest launch. Astra 6 was quickest overall, as it was in September. Speed was not the variable that decided anything here.

The verdict: 10–9–4

InterchangeFairgroundLaunchTotal
Claude Fable 5.124410
GPT-6 Astra4329
Grok 4.72114

Four points for a round win, scored on the finished artefact. Fable 5.1 and Astra 6 carry the scores from their September head-to-head, which used these same three briefs; Grok 4.7 is scored into that set. Scoring is done at BitsMinds from a side-by-side page with all the builds running at once. There is no blind scoring and no automated grading.

What ties Grok’s three rounds together

Read its faults next to each other and they stop looking like three unrelated slips.

On the interchange, the ramp geometry is correct and the road is painted over the top of it. On the fairground, three rides each consist of a core that rotates correctly inside a shell that was never attached to it. On the launch, the flight model is the most complete in the series and the camera is not connected to the vehicle it is following. In every case the mechanism is right and the layer the viewer actually sees is not joined to it.

That is a different failure from the ones this series usually finds. Astra’s carousel was drawn and never wired to turn; Sol’s ghosts had arcade-accurate rules and a one-line asymmetry that drove them through walls. Those are mechanisms that were missing or broken. Grok 4.7’s mechanisms are neither. They are built, they are correct, and they are running underneath something that does not reflect them — which is the one class of bug that looks completely fine in the source and completely wrong on the screen.

It is also, precisely, the class of bug that disappears the moment you open what you built. Every one of these takes seconds to see and none takes any thought to find: a line that does not open, a hoop that does not turn, a rocket that is not there. This series has spent since June arguing that the ban on self-checking is what separates these models, and the round that lifted the ban found both entrants verifying at length and still shipping a visible fault, because they verified the engine and never the level. Grok 4.7 has now produced the cleanest illustration of the same gap: it checked nothing, and everything it got wrong is something it would have seen.

All nine builds above are exactly as they came out, nothing added and nothing tidied. The interchanges reward a close look at where the ramps leave the road, the fairgrounds reward about fifteen seconds on each carousel, and the launches reward watching the clock at T+35.

More on Grok

Evergreen coverage we keep current — start here.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles