Part of our AI Comparisons
Models·12 min read
By BitsMinds Hands-on test

GPT-6 Sol vs GPT-5.6 Sol: A Small Step for Half the Price

After a poor Hill Climb result, is OpenAI’s GPT-6 Sol a step back from the GPT-5.6 Sol it replaces? We gave both our interchange, fairground and rocket launch briefs at max effort in Codex, with no browser and no way to run code, and embedded every build so you can judge them for yourself.

GPT-5.6 Sol and GPT-6 Sol: a small step at half the token price. Two equally sized metallic suns illuminate one miniature workbench with an interchange, a fairground and a rocket launch. Original BitsMinds editorial artwork for gpt-6-sol-vs-gpt-5-6-sol-build-off-2026. The half-price statement concerns the API input and output per-token list rates described in the article, not total run cost. Equal-size celestial instruments represent two Sol generations and do not encode scores. The models on the workbench symbolize the shared briefs rather than reproduce either delivered build. Official OpenAI outline preserved verbatim. Static, self-contained vector artwork. GPT-5.6 Sol GPT-6 Sol vs BITSMINDS LAB A SMALL STEP. HALF THE TOKEN PRICE. Three briefs. Two generations. BITSMINDS.COM
Share:
BitsMinds LabHow this test was run2 models · 3 briefs · one attempt per model per brief, except gpt-5 · Both at max in OpenAI Codex, not ultra, which adds automatic task delegation. GPT-5.6 Sol had built the interchange and the fairground in September at xhigh, then the top of the dial, and ran both again at max for this testShow details
Models
GPT-6 Sol (OpenAI Codex) · GPT-5.6 Sol (OpenAI Codex)
Attempts
One attempt per model per brief, except GPT-5.6 Sol's interchange, which was run twice: the first file drew every vehicle at the size of the whole scene, a slip the second run on the same prompt did not repeat. The second file is the one scored. No follow-up notes and no cleanup.
Reasoning effort
Both at max in OpenAI Codex, not ultra, which adds automatic task delegation. GPT-5.6 Sol had built the interchange and the fairground in September at xhigh, then the top of the dial, and ran both again at max for this test.
Tools allowed
Both ran in OpenAI Codex with the same ban on opening a browser or running code. GPT-5.6 Sol's run logs show nothing but file writes.
Timing
Wall-clock time from prompt to finished file: GPT-6 Sol 4m 45s, 6m 18s and 6m 36s on 22 September; GPT-5.6 Sol 6m 45s (second interchange run), 7m 15s and 8m 34s on 24 September.
Scoring
Points set round by round on the finished artefact: 2 and 1 on the interchange and the fairground, 2 each on the launch. GPT-6 Sol's builds are its 22 September files, scored again alongside GPT-5.6 Sol's. Final: GPT-6 Sol wins with 6 points, GPT-5.6 Sol takes 4.
Judging
Scored by BitsMinds from a local side-by-side comparison page of the untouched outputs. No blind scoring, no automated grading.
Published
September 24, 2026

Result

ModelProviderScore
GPT-6 SolOpenAI6 pts
GPT-5.6 SolOpenAI4 pts

GPT-6 Sol wins with 6 points, GPT-5.6 Sol takes 4. No regression on these briefs, and only a small, uneven step forward, from a model that was quicker on every brief and costs half as much per token.

The briefs — as described in the article; the exact prompt files were not published

  1. Round 1: Motorway interchange (SVG)

    A top-down view of a busy interchange as one self-contained SVG. Two roads crossing on different levels, connecting ramps, at least twelve vehicles following the curve of the road they are on, and correct layering so traffic on the lower road is hidden by the overpass.

    Result: GPT-6 Sol wins with 2 points, GPT-5.6 Sol takes 1. Both builds have ramps that go nowhere; GPT-5.6 Sol's loops also meet the road facing the wrong way, so its cars turn round on the spot, and its loop traffic vanishes under the flyover.

  2. Round 2: Seaside fairground (SVG)

    A seaside fairground at dusk, one SVG. A Ferris wheel with at least ten gondolas turning continuously, a spinning carousel with at least six horses rising and falling out of phase, twinkling bulbs, a reflection in the water. Each gondola must stay level with the horizon throughout the rotation.

    Result: GPT-6 Sol wins with 2 points, GPT-5.6 Sol takes 1. GPT-6 Sol's carousel turns with horses that never face the way they travel; GPT-5.6 Sol's carousel does not turn at all.

  3. Round 3: Rocket launch (web page)

    A self-contained web page that animates a rocket launching — countdown, ignition, liftoff and ascent — with a mission-control HUD. No library named, no 3D requested, one aesthetic instruction: make it look good.

    Result: A draw, 2 points each. Both flights stop at ascent; GPT-6 Sol frames its rocket too tightly, GPT-5.6 Sol's rocket is plainer and shivers on its own at liftoff but has the better framing and ending.

Read with care

  • Single run per model per brief, apart from GPT-5.6 Sol's second interchange run — a re-run could land differently.
  • Scored on the finished artefact only. Wall-clock time is reported but carries no points, and it depends on server load as well as the model: the two models ran on different days.
  • GPT-6 Sol's builds were made on 22 September and are scored again here rather than re-run.

Part of BitsMinds Lab, our series of original hands-on tests.

When GPT-6 Sol took on the Hill Climb brief in September, it built a shorter and easier game than GPT-5.6 Sol had, and that round called the step back plain. The names hide a change of rank. GPT-5.6 Sol was the flagship of its family; GPT-6 Sol sits in the middle of the GPT-6 range, below Astra, at $2 and $10 per million tokens against the $4 and $20 GPT-5.6 Sol still sells for. So is the new Sol a downgrade with a smaller price tag, or was Hill Climb one bad draw?

This test puts the two Sols side by side on the three briefs the series uses most: a motorway interchange seen from above and a seaside fairground at dusk, each as one animated SVG, and a cinematic rocket launch as one web page. GPT-6 Sol’s builds are the files it produced on 22 September for the GPT-6 Astra vs Sol vs Luna round. GPT-5.6 Sol ran on 24 September in OpenAI’s Codex on the same prompts, word for word, apart from the path each file was saved to. Both models ran at max reasoning effort, with no browser, no preview and no way to run code, so neither could look at what it had made.

GPT-5.6 Sol has met two of these briefs before. It built the interchange and the fairground in September for its round against Claude Opus 5, at xhigh, which was then the top of Codex’s dial. Codex has since added max, the setting GPT-6 Sol ran at, so both briefs were run again at max to keep the two models level. The rocket launch is GPT-5.6 Sol’s first.

GPT-6 Sol won two rounds and drew the third, and finished with 6 points to GPT-5.6 Sol’s 4.

Round 1: The interchange

The brief asks for a busy interchange as one self-contained SVG, with no text and no script. Two roads cross on different levels, ramps connect them, at least twelve vehicles follow the curves of the roads they drive on in both directions, and the layering hides the lower road’s traffic under the overpass. Two traps go unstated: whether each ramp really joins the roads at both of its ends, and whether traffic keeps to the right side after it merges.

GPT-5.6 Sol’s interchange was generated twice. In the first file every vehicle filled the whole frame. The cars are SVG symbols placed without a width or a height, and SVG then sizes each one to the full scene, so the picture is buried under 22 giant car bodies. That run’s log was clean and the file was complete. A second run on the same prompt did not repeat the slip, and the second file is the one scored and shown here, because it is the build a reader repeating the test would most likely get.

Animated motorway interchange built in one attempt by GPT-6 Sol
GPT-6 Sol, 4 minutes 45 seconds. Four flyover ramps that begin as stubs in the middle of the lower road.
Animated motorway interchange built by GPT-5.6 Sol
GPT-5.6 Sol, 6 minutes 45 seconds, second run. Four loops that meet the road facing the wrong way.

GPT-6 Sol built its interchange in 4 minutes 45 seconds. Twenty-four vehicles, sedans, vans and trucks with headlamps and red tail lights, drive on the ground-level east–west highway, across the north–south overpass and along four single-lane ramps. The ground is landscaped with planted verges, clumps of trees and drainage lines. The ramps are raised: each casts a shadow, the pairs that cross are drawn one above the other, and the east–west traffic passes beneath them. It is their ends that fail. At the lower road each ramp starts as a rounded stub sitting in the middle of the carriageway rather than peeling off a lane, and at the other end the ramp traffic is drawn underneath the overpass deck, so cars heading up to the bridge disappear beneath it instead of joining its lanes. Cars ride down the middle of a ramp and arrive nowhere. But there is less of it than in GPT-5.6 Sol’s build, and none of its ramps ends against the flow of traffic.

GPT-5.6 Sol’s second interchange took 6 minutes 45 seconds and looks the more ambitious of the two. A divided east–west highway runs at ground level, a north–south flyover sweeps across it in a long S-curve, and four cloverleaf loops link them. Around the junction sit small industrial parcels, a drainage channel, service roads, clumps of trees, lamp posts with a slow glow, and two green road signs with arrows. Both roads keep left, consistently, like a British motorway. Its 24 vehicles all enter and leave off-screen, so no car appears from nowhere at the edge of the frame. The loops are where it falls apart. Each loop meets its road facing the wrong way, so wherever a loop and a road join, the car turns round on the spot: on the north-west loop, a car heading east along the lower road reaches the loop’s mouth, snaps round 171 degrees and heads back west into it. Every loop route also runs along the flyover for part of its length, and there its cars are drawn underneath the flyover’s road surface, where they vanish. Sampled across a loop, a car spends 18 to 37 per cent of its time on screen hidden beneath the flyover. Cars pass over the junction, under it, and into it without coming out.

Round 1 goes to GPT-6 Sol, which takes 2 points; GPT-5.6 Sol takes 1. Both builds have ramps that go nowhere. GPT-6 Sol’s make less of a mess.

Round 2: The fairground

A seaside fairground at dusk, one SVG: a Ferris wheel with at least ten gondolas turning continuously and staying level with the horizon all the way round, a carousel with at least six horses rising and falling out of phase under a spinning canopy, twinkling bulbs, lights chasing round the rim of the wheel, a low sun or moon, and the whole scene reflected in the water. The gondolas are the trap the brief states. The carousel is the one it does not.

Animated seaside fairground at dusk built in one attempt by GPT-6 Sol
GPT-6 Sol, 6 minutes 18 seconds. A carousel that turns, horses that never face the way they travel, and water with no reflection in it.
Animated seaside fairground at dusk built in one attempt by GPT-5.6 Sol
GPT-5.6 Sol, 7 minutes 15 seconds. A correct Ferris wheel beside a carousel that stands still.

GPT-6 Sol’s scene, built in 6 minutes 18 seconds, is a pier at sunset under a band of clouds, with a low sun, a headland in the distance, a ticket kiosk, a “SWEETS” taffy booth, lamp posts, bunting and two long festoons of bulbs that twinkle independently. Its wheel stands on an A-frame and carries twelve cabins. Rim, spokes and cabins turn together once every 48 seconds while each cabin counter-rotates and sways, so the stated trap is cleared, and two rings of chasing lights run round the rim. Its carousel does turn. A striped, sloping canopy rotates every twelve seconds, and eight horses ride an elliptical track at the same speed, each on its own pole, rising and falling at staggered offsets. What they never do is turn: every horse points the same way all the way round the track, so on the far side of the ride it travels backwards. It wrote a reflection too, then clipped it with a rectangle placed in the flipped coordinates, which cuts every reflected ride away. The water shows ripples and nothing else.

GPT-5.6 Sol took 7 minutes 15 seconds over a purple dusk with a setting sun and a path of light across the sea, a ticket booth topped with a star, a striped tent, a boardwalk with railings and festoons of bulbs strung across the sky. Its Ferris wheel is correct: twelve gondolas stay level while the wheel turns once every 24 seconds, each swaying a few degrees either way, and the rim lights travel round with the wheel. Its reflection is there as well, a flattened, rippling copy of the whole fair. The carousel does not turn at all. Its only motion is a dashed stroke running along the edge of the canopy and a string of chasing bulbs below it. The canopy, the poles and the seven horses stay put, and the horses bob on the spot, ten pixels up and down, without travelling anywhere; the first and the last of them bob in lockstep. The source leaves no doubt about the intent, with a comment that reads “Poles stay fixed while seven horses move independently”. This is a step back from its own September fairground, whose canopy at least turned while the horses bobbed in place. Across the scene the drawing is also plainer than GPT-6 Sol’s.

Round 2 goes to GPT-6 Sol, which takes 2 points; GPT-5.6 Sol takes 1.

Round 3: The rocket launch

The only open brief. It names no library, asks for no 3D, and its one instruction on style is to make it look good. It wants a countdown, ignition, liftoff and a sustained ascent, with a mission-control HUD reporting the state of the flight.

AURORA 06 — GPT-6 Sol, 6 minutes 36 seconds.

ORBITAL // 56 — GPT-5.6 Sol, 8 minutes 34 seconds. Both start on their own countdown.

GPT-6 Sol’s AURORA 06 stands a three-core rocket beside its service tower against dark mountains, under a headline that reads “The sky awaits”. A ten-second count leads into three seconds of ignition, with flame and smoke building under the pad before the rocket leaves the tower. The HUD reports altitude, velocity and downrange distance, lights up a four-step phase track and runs a mission clock. Once the vehicle clears the tower the camera holds it at a fixed height in the frame while the ground falls away, and during the climb the rocket’s nose sits cut off at the top of the frame: the shot is too tight. The flight never goes past “ascent”. The engines stay at 100%, nothing separates, and the climb carries on for as long as the page stays open.

GPT-5.6 Sol’s ORBITAL // 56 has the fuller mission-control desk: a vehicle-systems panel, a telemetry block with altitude, velocity, throttle, g-load, downrange distance and dynamic pressure, a fuel gauge, and a timeline of the flight’s milestones along the bottom of the screen. The count runs from T−10 through strongback retraction and launch commit to ignition, and a three-second hold-down on the pad before liftoff. The rocket itself is less detailed than GPT-6 Sol’s, and its launch has an odd shiver. From ignition the vehicle alone jitters sideways, several pixels at a time and several times a second, peaking at liftoff and fading over the first eleven seconds of flight, while the tower and the ground stay perfectly still around it. The framing is better. The camera keeps the whole vehicle in shot, nose to exhaust, and the ending is the prettier of the two: as the rocket leaves the atmosphere the stars start to streak past it, which is more dramatic than it is realistic. It stops where GPT-6 Sol stops. The throttle holds at 103%, the altitude stops climbing at 610 km and the velocity at 7.82 km/s, the fuel never runs out and the side boosters never separate.

Neither model pulled in a library, though the brief allows any from a public CDN. Counting every model that has taken this brief since September, that makes nine models from three labs drawing their launches by hand.

Round 3 is a draw: GPT-6 Sol and GPT-5.6 Sol take 2 points each.

Quicker, and cheaper to run

BuildGPT-6 SolGPT-5.6 Sol
Interchange4m 45s6m 45s (second run)
Fairground6m 18s7m 15s
Rocket launch6m 36s8m 34s
All three17m 39s22m 34s
Output tokens, all three55,90272,666
At API list pricesabout $0.76about $1.90

GPT-6 Sol was the quicker model on every brief, and it wrote less to get there: 55,902 output tokens across the three builds against 72,666, by Codex’s own count. Priced at each model’s API list rate from the token counts Codex reports, its three builds come to about 76 cents and GPT-5.6 Sol’s to about $1.90, two and a half times as much. The runs themselves went through a ChatGPT subscription, so no one was billed those amounts; they show how the two price tags play out on the same work. Artificial Analysis finds a similar gap at larger scale. On its Intelligence Index, v4.3.2 as captured on 22 September, GPT-6 Sol scores 48 against GPT-5.6 Sol’s 47 and costs $1.06 per task against $1.99, both at max, as our leaderboard shows.

The conditions were the same where it mattered: the same prompts word for word, one model per column, max effort, no browser and no code execution. Two things differ. GPT-5.6 Sol’s interchange had the second run described above. And wall-clock time depends on the server as well as the model: GPT-6 Sol’s three runs went out on 22 September, GPT-5.6 Sol’s on 24 September. The Codex logs for all of GPT-5.6 Sol’s runs show nothing but file writes; it did not search Codex’s own memory before starting, as it did in the second Hill Climb round.

The verdict: GPT-6 Sol wins with 6 points, GPT-5.6 Sol takes 4

InterchangeFairgroundLaunchTotal
GPT-6 Sol2226
GPT-5.6 Sol1124

Points were set round by round on the finished artefact alone: 2 and 1 on the interchange and the fairground, 2 each on the launch. GPT-6 Sol’s builds are its 22 September files, scored again next to GPT-5.6 Sol’s, which is why its interchange scores differently here than it did against Astra. The builds were scored at BitsMinds from a local side-by-side page of the untouched outputs, with no blind scoring and no automated grading. Our GPT-6 Sol review weighs these builds against the independent benchmark numbers.

No regression, and not much progress

On these three briefs the new Sol is not a step back. It is at least as good as the model it replaces in every round and better in two, and that is with GPT-5.6 Sol given a second attempt at the interchange. It is not a large step forward either. Both interchanges send cars nowhere, both flights end at “ascent”, and GPT-5.6 Sol still framed its rocket better and gave it the better ending. The improvement is small, and it does not show up everywhere.

What changes the sum is the price. GPT-6 Sol does slightly better work for half the per-token rate, and on these builds it used fewer tokens too. A model that holds its predecessor’s level, edges past it in places and costs about half as much to run is a good deal, whatever Hill Climb said about it. Hill Climb still says it: the same model built a short, forgiving game there. On a single brief, a model can land well below its usual level.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Hill Climb build-off: Opus 5.5's red buggy jumps toward a trail of gold coins, with GPT-6 Sol's orange 06 hatchback and Grok 4.7's teal 47 buggy on the surrounding hills. Original vector editorial illustration for the BitsMinds Hill Climb comparison. Vehicle designs are informed by screenshots of Hill Hopper, Ridge Runner and Switchback. The scene combines the separate games as an editorial montage, not a screenshot or a shared in-game race. No performance scores are encoded. Official provider logo outlines are preserved verbatim. 06 47 BITSMINDS LAB / HILL CLIMB FINALLY, A REAL GAME. BITSMINDS.COM Opus 5.5 vs GPT-6 Sol vs Grok 4.7
Models

Opus 5.5 vs GPT-6 Sol vs Grok 4.7: Finally, a Real Game

OPEN WEIGHTS · MIT LICENCE #1 OPEN-WEIGHT MODEL BITSMINDS.COM
Models

Xiaomi's MiMo-V2.6 Pro Is Now the Top Open-Weight Model

GPT-6 family build-off: Astra, Sol and Luna, represented by a crystal star, a golden sun and a silver crescent, on one silver astronomical instrument against a midnight-blue sky. Original vector editorial illustration for BitsMinds. The three equal-height instruments symbolize the GPT-6 family. Their shared base represents the shared build briefs; dimensions and celestial forms do not encode scores or performance. OpenAI logo outline preserved verbatim from the project asset. Static, self-contained SVG with no raster, filters, scripts or external dependencies. GPT-6 THE FAMILY BUILD-OFF OpenAI Astra Sol Luna VS VS BITSMINDS LAB SAME BRIEFS. THREE MODELS. BITSMINDS.COM
Models

GPT-6 Astra vs Sol vs Luna: Is the Discount Worth It?