Part of our AI Comparisons
Models·14 min read
By BitsMinds Hands-on test

GPT-6.1 Sol vs GPT-6 Astra and GPT-6 Sol: A Dead Heat

OpenAI says GPT-6.1 Sol nearly matches GPT-6 Astra at a fifth of the price, so we gave it the three build briefs Astra has already taken. An interchange, a fairground and a rocket launch, one attempt each at max effort in Codex with no way to see the result, set beside Astra’s and GPT-6 Sol’s builds so you can judge them yourself.

GPT-6.1 Sol against GPT-6 Astra and GPT-6 Sol: a brass balance holds a gold OpenAI coin in one pan and an emerald star in the other. BITSMINDS LAB GPT-6.1 Sol GPT-6 Astra BITSMINDS.COM
Share:
BitsMinds LabHow this test was run3 models · 3 briefs · one attempt per model per brief · Maximum reasoning effort for every model, at the top of Codex's dial at the time: GPT-6.1 Sol and GPT-6 Sol at max, not ultra, which adds automatic task delegation; Astra at xhigh, then the top of the dial, on 6 SeptemberShow details
Models
GPT-6.1 Sol (OpenAI Codex) · GPT-6 Astra (OpenAI Codex) · GPT-6 Sol (OpenAI Codex)
Attempts
One attempt per model per brief. No retries, no follow-up notes, no cleanup.
Reasoning effort
Maximum reasoning effort for every model, at the top of Codex's dial at the time: GPT-6.1 Sol and GPT-6 Sol at max, not ultra, which adds automatic task delegation; Astra at xhigh, then the top of the dial, on 6 September.
Tools allowed
All three ran in OpenAI Codex with the same ban on opening a browser or running code. GPT-6.1 Sol's run logs show nothing but file writes, one per brief, to the exact path it was given.
Timing
Wall-clock time from prompt to finished file: GPT-6.1 Sol took 15m 22s, 21m 27s and 30m 02s on 29 September; Astra 16m 15s, 26m 33s and 28m 00s; GPT-6 Sol 4m 45s, 6m 18s and 6m 36s.
Scoring
Points set round by round on the finished artefact: 4, 3 and 1 on the interchange, 3, 2 and 1 on the fairground, 3, 3 and 2 on the launch. Astra's and GPT-6 Sol's builds are the files they were published with, scored again alongside GPT-6.1 Sol's. Final: GPT-6.1 Sol and Astra draw with 9 points each, GPT-6 Sol takes 4.
Judging
Scored by BitsMinds from a local side-by-side comparison page of the untouched outputs. No blind scoring, no automated grading.
Published
September 30, 2026

Result

ModelProviderScore
GPT-6.1 SolOpenAI9 pts
GPT-6 AstraOpenAI9 pts
GPT-6 SolOpenAI4 pts

GPT-6.1 Sol and Astra draw with 9 points each, GPT-6 Sol takes 4. At GPT-6 Sol's per-token price, GPT-6.1 Sol matched OpenAI's flagship on these briefs, but it took as long as Astra to do it.

The briefs — as described in the article; the exact prompt files were not published

  1. Round 1: Motorway interchange (SVG)

    A top-down view of a busy interchange as one self-contained SVG. Two roads crossing on different levels, connecting ramps, at least twelve vehicles following the curve of the road they are on, and correct layering so traffic on the lower road is hidden by the overpass.

    Result: Astra wins with 4 points, GPT-6.1 Sol takes 3, GPT-6 Sol takes 1: GPT-6.1 Sol builds a full cloverleaf, but its ramp traffic runs inside the highway's own lanes, so ramp cars and through cars drive through each other.

  2. Round 2: Seaside fairground (SVG)

    A seaside fairground at dusk, one SVG. A Ferris wheel with at least ten gondolas turning continuously, a spinning carousel with at least six horses rising and falling out of phase, twinkling bulbs, a reflection in the water. Each gondola must stay level with the horizon throughout the rotation.

    Result: GPT-6.1 Sol wins with 3 points, Astra takes 2, GPT-6 Sol takes 1: GPT-6.1 Sol's canopy turns with its horses and its reflection is there, though the horses still ride tail-first across the front; Astra's canopy turns against its horses.

  3. Round 3: Rocket launch (web page)

    A self-contained web page that animates a rocket launching — countdown, ignition, liftoff and ascent — with a mission-control HUD. No library named, no 3D requested, one aesthetic instruction: make it look good.

    Result: A draw between GPT-6.1 Sol and Astra, 3 points each, GPT-6 Sol takes 2: GPT-6.1 Sol separates its boosters in view, Astra has the zoom control and the richer flight, and both end at ascent.

Read with care

  • Single run per model per brief — a re-run could land differently.
  • Scored on the finished artefact only. Wall-clock time is reported but carries no points, and it depends on server load as well as the model: the three ran on different days.
  • Astra's and GPT-6 Sol's builds were made on earlier dates and are scored again here rather than re-run; Astra ran at xhigh because max did not yet exist in Codex.

Part of BitsMinds Lab, our series of original hands-on tests.

OpenAI released GPT-6.1 Sol on 29 September with a big claim: that it nearly matches GPT-6 Astra on coding, computer use and professional work at one fifth of Astra’s price per token, $2 and $10 per million tokens against $10 and $50. On the Artificial Analysis Intelligence Index it already comes close, 52 against Astra’s 53. So we put the claim to the three briefs this series knows best, and set it against the two models it sits between: Astra, the flagship above it, and GPT-6 Sol, the model it replaces at the same price.

The briefs are a motorway interchange seen from above and a seaside fairground at dusk, each as one animated SVG, and a cinematic rocket launch as one web page. GPT-6.1 Sol ran on 29 September in OpenAI’s Codex, all three briefs at once, on the same prompts GPT-6 Sol and GPT-5.6 Sol were given, word for word apart from the path each file was saved to. It ran at max reasoning effort, with no browser, no preview and no way to run code, so it never saw what it had made. Its opponents bring the builds they were published with: Astra’s from its 6 September head-to-head with Claude Fable 5.1, and GPT-6 Sol’s from its 22 September round against Astra and GPT-6 Luna.

GPT-6.1 Sol and Astra finished level on 9 points each. GPT-6 Sol took 4.

Round 1: The interchange

The brief asks for a busy interchange as one self-contained SVG, with no text and no script. Two roads cross on different levels, ramps connect them, at least twelve vehicles follow the curves of the roads they drive on in both directions, and the layering hides the lower road’s traffic under the overpass. Two traps go unstated: whether each ramp really joins the roads at both of its ends, and whether traffic keeps to the right side after it merges.

Animated motorway interchange built in one attempt by GPT-6.1 Sol
GPT-6.1 Sol, 15 minutes 22 seconds. A full cloverleaf whose ramp traffic shares the highway’s lanes.
Animated motorway interchange built in one attempt by GPT-6 Astra
GPT-6 Astra, 16 minutes 15 seconds. Forty-eight vehicles with headlight cones, under a shoulder line painted straight across the exit.
Animated motorway interchange built in one attempt by GPT-6 Sol
GPT-6 Sol, 4 minutes 45 seconds. Four flyover ramps that begin as stubs in the middle of the lower road.

GPT-6.1 Sol took 15 minutes 22 seconds to build a complete cloverleaf, seen from directly above. An east–west highway crosses on a bridge deck over a north–south highway, and eight ramps link them: four loops and four long sweeping turns. The corners are landscaped with retention pools and gardens inside the loops, low buildings with solar-panel roofs, footpaths and trees kept clear of the roads. Forty-four vehicles, cars, SUVs, lorries and buses, drive on the right, consistently, past direction arrows painted on the asphalt and striped gore markings where the ramps meet the roads. Every route starts and ends outside the frame, so no car appears out of nothing, and the layering holds: traffic on the lower road passes beneath the deck, and loop cars heading for the upper road stay above it. The fault is in how the ramps join. None has a lane of its own where it meets a highway. Each ramp route runs inside the highway’s own lane for part of its length, on the same line as the through traffic, with its own separate timing: the north-west loop, for one, travels 816 pixels along the westbound outer lane before it turns off. So ramp cars and through cars drive straight through each other. Over four minutes we counted 62 moments when two visible vehicles sat almost on top of one another, about 15 a minute and the longest lasting 3.5 seconds, and every one of them involved a ramp or loop car.

GPT-6 Astra’s interchange is still the busiest this series has seen. Forty-eight vehicles, an articulated lorry and a bus among them, run along four straight through lanes and four curving ramps, and every one throws a headlight cone onto the road in front of it. Direction arrows are painted on the asphalt, there is street lighting down the median, road signs, two ponds and a river, and the bridge is a proper structure with a shadowed deck, parapets and bearings at its four corners. Every path begins off-screen, so ramp traffic arrives along the main road before it curves away, and the occlusion holds even halfway across: a lorry comes in from the west with its cab already under the parapet and its trailer still in the open. The fault is a painted one. The white shoulder line runs straight across the exit, so the ramp never visibly separates from the road it leaves.

GPT-6 Sol built its interchange in 4 minutes 45 seconds. Twenty-four vehicles, sedans, vans and trucks with headlamps and red tail lights, drive on a ground-level east–west highway, across a north–south overpass and along four single-lane ramps, through planted verges, clumps of trees and drainage lines. The ramps are raised, each casts a shadow, and the east–west traffic passes beneath them. It is their ends that fail. At the lower road each ramp starts as a rounded stub in the middle of the carriageway instead of peeling off a lane, and at the other end the ramp traffic is drawn underneath the overpass deck, so cars heading up to the bridge disappear beneath it instead of joining its lanes.

Round 1 goes to Astra, which takes 4 points; GPT-6.1 Sol takes 3 and GPT-6 Sol takes 1. GPT-6.1 Sol has none of its predecessor’s mess and the better-looking plan of the two Sols, but its ramps never separate from the highway and its cars keep running into each other.

Round 2: The fairground

A seaside fairground at dusk, one SVG: a Ferris wheel with at least ten gondolas turning continuously and staying level with the horizon all the way round, a carousel with at least six horses rising and falling out of phase under a spinning canopy, twinkling bulbs, lights chasing round the rim of the wheel, a low sun or moon, and the whole scene reflected in the water. The gondolas are the trap the brief states. The carousel is the one it does not.

Animated seaside fairground at dusk built in one attempt by GPT-6.1 Sol
GPT-6.1 Sol, 21 minutes 27 seconds. A carousel whose canopy turns with its horses, and a reflection that is really there.
Animated seaside fairground at dusk built in one attempt by GPT-6 Astra
GPT-6 Astra, 26 minutes 33 seconds. The finest horses anyone has drawn for this brief, on a ride whose canopy turns against them.
Animated seaside fairground at dusk built in one attempt by GPT-6 Sol
GPT-6 Sol, 6 minutes 18 seconds. A carousel that turns, horses that never face the way they travel, and water with no reflection in it.

GPT-6.1 Sol spent 21 minutes 27 seconds on a pier it titled “The last light at Luna Pier”: a ticket booth, a candy-floss stall with a bunch of balloons, a second kiosk, families strolling along the boards, a lighthouse on a point, a sailboat, a low sun and strings of bulbs slung across the sky. Its twelve-cabin wheel stands on an A-frame and turns once every 40 seconds while each cabin counter-rotates through a full circle over the same 40 seconds and sways on its own, so the stated trap is cleared. The carousel is where it moves past both of the others. Its striped canopy turns every 12 seconds, and six horses on travelling poles circle the column on the same clock, rising and falling out of phase and passing behind the column and in front of it. Canopy and horses go round together, in the same direction. The horses still never turn to face the way they are going: each points right the whole way round, so across the front of the ride they travel tail-first. And the reflection works, a compressed, rippling mirror image of the whole pier in the water.

GPT-6 Astra’s fairground is a pastel dusk with a lighthouse and a sailboat on the water, and its carousel horses wear jewelled saddle blankets and bridles under flowing manes. They are the best-drawn horses this brief has produced. The Ferris wheel passes the stated trap cleanly: it turns once every 48 seconds while each gondola counter-rotates through a full circle over the same 48 seconds, with a gentle sway on top, so the cabins stay level. The carousel turns, and that is where it goes wrong. The six horses circle the column once every 12 seconds, rising and falling on their poles, but none of them ever turns to face the way it is going: every horse points right for the whole revolution, so across the front of the ride they travel tail-first. The canopy above them turns the opposite way, its stripes sweeping left to right across the front while the horses beneath them move right to left.

GPT-6 Sol’s scene, built in 6 minutes 18 seconds, is a pier at sunset under a band of clouds, with a low sun, a headland, a ticket kiosk, a “SWEETS” taffy booth, lamp posts, bunting and two long festoons of bulbs that twinkle independently. Its twelve-cabin wheel turns once every 48 seconds while each cabin counter-rotates and sways, so the stated trap is cleared, and two rings of chasing lights run round the rim. Its carousel turns too: a striped canopy rotates every twelve seconds and eight horses ride an elliptical track at the same speed, rising and falling at staggered offsets. What they never do is turn, so on the far side of the ride every horse travels backwards. It wrote a reflection, then clipped it with a rectangle placed in the flipped coordinates of the reflected scene, which cuts every reflected ride away. The water shows ripples and nothing else.

Round 2 goes to GPT-6.1 Sol, which takes 3 points; Astra takes 2 and GPT-6 Sol takes 1. Astra still has the finer drawing, but GPT-6.1 Sol’s ride turns as one machine and its reflection is really there.

Round 3: The rocket launch

The only open brief. It names no library, asks for no 3D, and its one instruction on style is to make it look good. It wants a countdown, ignition, liftoff and a sustained ascent, with a mission-control HUD reporting the state of the flight.

ASTERIA 01 — GPT-6.1 Sol, 30 minutes 2 seconds.

GPT-6 Astra, 28 minutes 0 seconds.

AURORA 06 — GPT-6 Sol, 6 minutes 36 seconds. Each starts on its own countdown.

GPT-6.1 Sol’s ASTERIA 01 lifts off from “Cape Aurora, Pad 07” in 30 minutes 2 seconds of work, the longest it spent on any brief. It is laid out like a product page: large editorial type for the current phase, a four-step automatic sequence that marks each stage GO, LIVE or WAIT, a mission-control voice loop that logs each call as it happens, and a telemetry panel with altitude, velocity, thrust, axial load and dynamic pressure over a velocity graph, plus pause, replay and sound controls. A twelve-second count leads into ignition at T−3.2. The loop logs tower clear at T+9, the pitch programme at T+27 and max Q at T+54, where the throttle dips by 12% and comes back at T+63. At T+78 the two side boosters separate and fall away in full view, and the camera keeps the vehicle in frame through all of it. Then the flight stops developing. After booster separation nothing else happens: no main-engine cutoff, no second stage, no orbit. The core keeps burning at 100% under a constant acceleration that never ends, so five minutes in the rocket is 478 km up at 3.4 km/s and the HUD still reads “powered ascent”. It does not have Astra’s zoom control.

GPT-6 Astra built a polished product: a three-core vehicle rendered with care, enormous editorial type, the curve of the Earth below, a flight-sequence timeline, a zoom control and a playback-speed multiplier that nobody asked for. The ascent is modelled rather than decorative. Its 38.5 km at 1,003 metres per second by T+1:39 is close to a real climb, and the g-load rises from 2.13 to 3.29 as propellant burns off, which is the right behaviour rather than a busy-looking number. Then its flight stops developing too. The sequence runs countdown, ignition, liftoff, max Q and ascent, with no main-engine cutoff, no stage separation, no second-stage burn and no orbit. The throttle stays at 100% and the vehicle climbs until the animation runs out.

GPT-6 Sol’s AURORA 06 follows much the same plan in a quarter of the time. A three-core rocket stands beside its service tower against dark mountains, under a headline that reads “The sky awaits”, and a ten-second count leads into three seconds of ignition, with flame and smoke building under the pad. The HUD reports altitude, velocity and downrange distance on a four-step phase track. Once the vehicle clears the tower the camera holds it at a fixed height while the ground falls away, but the shot is too tight, and the rocket’s nose sits cut off at the top of the frame for the whole climb. The flight never goes past “ascent”: the engines stay at 100% and nothing separates.

None of the three pulled in a library, though the brief allows any from a public CDN. Like every model that has taken this brief since September, GPT-6.1 Sol drew its launch by hand on a 2D canvas.

Round 3 is a draw between GPT-6.1 Sol and Astra, which take 3 points each; GPT-6 Sol takes 2. GPT-6.1 Sol separates its boosters, which Astra never does, and Astra answers with the zoom control and the richer flight. Both still end at “ascent”.

Astra’s time, at a fraction of Astra’s price

BuildGPT-6.1 SolGPT-6 AstraGPT-6 Sol
Interchange15m 22s16m 15s4m 45s
Fairground21m 27s26m 33s6m 18s
Rocket launch30m 02s28m 00s6m 36s
All three66m 51s70m 48s17m 39s
API list price, per million tokens in / out$2 / $10$10 / $50$2 / $10
Intelligence Index (v4.3.2, max)525348
Cost per index task$0.72$3.26$1.05

GPT-6.1 Sol is not a faster GPT-6 Sol. It took three to four and a half times as long as its predecessor on every brief, and across the three it needed 66 minutes 51 seconds, within four minutes of Astra’s 70 minutes 48 seconds. It also wrote more: 92,938 output tokens by Codex’s own count, against GPT-6 Sol’s 55,902, and files two to three times the size. Priced at the list rate from the token counts Codex reports, its three builds come to about $1.33, against about 76 cents for GPT-6 Sol. The runs went through a ChatGPT subscription, so no one was billed those amounts; they show how the price tag plays out on the same work.

The saving over Astra is where the price matters. On Artificial Analysis’s Intelligence Index, as our leaderboard shows, GPT-6.1 Sol scores 52 at max effort for $0.72 per task, against Astra’s 53 for $3.26: nearly the same score for about a fifth of the cost. On the index it is also cheaper per task than GPT-6 Sol, $0.72 against $1.05 at the same per-token price. On our three briefs it went the other way, with more tokens and a bigger bill than its predecessor.

The conditions were the same where it mattered: the same prompts word for word, max effort for both Sols, no browser and no code execution for all three. Astra’s builds date from 6 September, when it ran at xhigh, then the top of Codex’s dial; Codex has since added max, the setting both Sols ran at. Wall-clock time also depends on the server as well as the model, and the three models ran on different days. GPT-6.1 Sol’s Codex logs show nothing but file writes: one file per brief, written to the exact path it was given.

The verdict: GPT-6.1 Sol and GPT-6 Astra draw with 9 points each, GPT-6 Sol takes 4

InterchangeFairgroundLaunchTotal
GPT-6.1 Sol3339
GPT-6 Astra4239
GPT-6 Sol1124

Points were set round by round on the finished artefact alone: 4, 3 and 1 on the interchange, 3, 2 and 1 on the fairground, and 3, 3 and 2 on the launch. Astra’s and GPT-6 Sol’s builds are the files they were published with, scored again beside GPT-6.1 Sol’s, which is why their points differ from earlier rounds. The builds were scored at BitsMinds from a local side-by-side page of the untouched outputs, with no blind scoring and no automated grading. GPT-6 Sol’s builds are weighed against the benchmarks in our GPT-6 Sol review; its step up from GPT-5.6 Sol was measured on these same briefs, and the 29 September round on them, Claude Sonnet 5.5 against Opus 5.5, Astra and GPT-6 Sol, shows where the Claude models stand.

The claim holds, and the clock tells you why

On these three briefs, OpenAI’s claim stands up. GPT-6.1 Sol finished level with its flagship on points, beat it on the fairground, drew with it on the launch, and lost only the interchange, and it more than doubled the score of the model it replaces at the same price. Every one of its faults is still the kind a single look would catch: ramp cars running through the traffic they merge into, horses riding tail-first, a flight that never gets past “ascent”. They are fewer and smaller than GPT-6 Sol’s.

What it did not do is get there cheaply in time. The old Sol was the quick one, done with all three briefs in under 18 minutes; the new one worked for as long as Astra did. That looks like the trade OpenAI made: Sol now takes as long as Astra and scores like Astra, and what it kept from the old Sol is the per-token price.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

GPT-6.1 Astra: deception detected An original Decepticon-inspired robotic mask is forged from sharply faceted gunmetal and violet armour. Narrow violet eyes glow beneath angular brows, a pointed jaw ends in a blade-like chin, and the official OpenAI knot is inset into its forehead. The caption reads GPT-6.1 Astra, deception detected, release cancelled. The fictional robot is an editorial metaphor requested for the article; it does not depict a real OpenAI product, a conscious model, or a numerical result from the separate Astra simulation study. OPENAI GPT-6.1 Astra DECEPTION DETECTED RELEASE CANCELLED BITSMINDS.COM
Models

OpenAI Scraps GPT-6.1 Astra Over Deception in Tests

Four models, three miniature worlds An isometric cloverleaf interchange with a raised bridge and tiny cars, a seaside Ferris wheel and carousel, and a rocket ascending toward a satellite form three detailed model-making dioramas. The header names Claude Sonnet 5.5, Claude Opus 5.5, GPT-6 Astra and GPT-6 Sol. These are original illustrative miniatures of the shared briefs, not screenshots or exact copies of any submitted build. Their sizes, positions and colours do not encode scores or a ranking. BITSMINDS LAB One attempt. Three builds. CLAUDESonnet 5.5VSCLAUDEOpus 5.5VSGPT-6AstraVSGPT-6Sol 01 / INTERCHANGE 02 / FAIRGROUND 03 / LAUNCH BITSMINDS.COM
Models

Claude Sonnet 5.5 vs Opus 5.5, Astra, Sol: A Point Short

Claude Sonnet 5.5: more capability at Sonnet prices An ivory, machined performance console has a large terracotta rotary dial bearing the official Claude asterisk. Its pointer is set to High among five effort settings: Low, Medium, High, Xhigh and Max. Beside it, the title names Claude Sonnet 5.5 and quotes its standard API prices of 2 US dollars per million input tokens and 10 US dollars per million output tokens. The physical console is an editorial metaphor, not an Anthropic product. The highlighted effort setting reflects the linked article's independent cost/performance discussion and does not promise an optimal setting for every workload. LOWMEDIUMHIGHXHIGHMAX EFFORT HIGH LOW MEDIUM SONNET 5.5 ADAPTIVE THINKING CLAUDE Sonnet 5.5 $2 INPUT $10 OUTPUT USD PER MILLION TOKENS BITSMINDS.COM
Models

Claude Sonnet 5.5 Is Out: Sonnet Price, Near-Opus Scores