GPT-6 Astra vs Sol vs Luna: Is the Discount Worth It?
OpenAI's two cheaper GPT-6 models took on the three briefs GPT-6 Astra answered in September, and the flagship won every round. Astra finished with 10 points, Sol took 5 and Luna took 3, the same order as the price list. Sol and Luna built each brief two to six times faster, and every one of their six builds shipped a fault you can see.
BitsMinds LabHow this test was run3 models · 3 briefs · one attempt per model per brief · GPT-6 Sol and GPT-6 Luna at max in OpenAI Codex, a level Codex added after Astra's September runs at xhigh, which was then the top of its dial. Not ultra, which adds automatic task delegationShow details
- Models
- GPT-6 Astra (OpenAI Codex) · GPT-6 Sol (OpenAI Codex) · GPT-6 Luna (OpenAI Codex)
- Attempts
- One attempt per model per brief. No retries, no follow-up notes, no cleanup.
- Reasoning effort
- GPT-6 Sol and GPT-6 Luna at max in OpenAI Codex, a level Codex added after Astra's September runs at xhigh, which was then the top of its dial. Not ultra, which adds automatic task delegation.
- Tools allowed
- All three ran in OpenAI Codex with the same ban on opening a browser or running code. The Sol and Luna run logs show nothing but file writes.
- Timing
- Wall-clock time from prompt to finished file: Sol 4m 45s, 6m 18s and 6m 36s; Luna 8m 06s, 9m 06s and 4m 29s; Astra's September runs 16m 15s, 26m 33s and 28m 00s. The six Sol and Luna runs went out in parallel on one evening.
- Scoring
- Points set round by round on the finished artefact: 4, 1 and 1 on the interchange, 3, 2 and 1 on the fairground and the launch. Astra's builds are its September files, scored again alongside Sol and Luna. Final: Astra wins with 10 points, Sol takes 5, Luna takes 3.
- Judging
- Scored by BitsMinds from a local side-by-side comparison page of the untouched outputs. No blind scoring, no automated grading.
- Published
- September 23, 2026
Result
| Model | Provider | Score |
|---|---|---|
| GPT-6 Astra | OpenAI | 10 pts |
| GPT-6 Sol | OpenAI | 5 pts |
| GPT-6 Luna | OpenAI | 3 pts |
Astra wins with 10 points, Sol takes 5, Luna takes 3, the same order as the price list. Sol and Luna finished each brief two to six times faster, and each of their six builds shipped a fault a single look would catch.
The briefs — as described in the article; the exact prompt files were not published
Round 1: Motorway interchange (SVG)
A top-down view of a busy interchange as one self-contained SVG. Two roads crossing on different levels, connecting ramps, at least twelve vehicles following the curve of the road they are on, and correct layering so traffic on the lower road is hidden by the overpass.
Result: Astra wins with 4 points; Sol and Luna take 1 each. Both new builds have ramps that start in the middle of a road and connect to nothing, their traffic hidden under the overpass deck; Astra's only fault is a shoulder line painted across its exit.
Round 2: Seaside fairground (SVG)
A seaside fairground at dusk, one SVG. A Ferris wheel with at least ten gondolas turning continuously, a spinning carousel with at least six horses rising and falling out of phase, twinkling bulbs, a reflection in the water. Each gondola must stay level with the horizon throughout the rotation.
Result: Astra wins with 3 points, Sol takes 2, Luna takes 1: Sol's carousel turns but its horses never face the way they travel and its reflection is clipped away; Luna's reflection works and its carousel spins like a second Ferris wheel.
Round 3: Rocket launch (web page)
A self-contained web page that animates a rocket launching — countdown, ignition, liftoff and ascent — with a mission-control HUD. No library named, no 3D requested, one aesthetic instruction: make it look good.
Result: Astra wins with 3 points, Sol takes 2, Luna takes 1: Sol's flight, like Astra's, never gets past ascent; Luna plans staging and orbit, but its camera offset is counted twice and the rocket leaves the frame by T+11.
Read with care
- Single run per model — a re-run could land differently.
- Scored on the finished artefact only. Wall-clock time is reported but carries no points, and it depends on server load as well as the model: Astra's runs were on other days.
- Sol and Luna ran one effort level higher than Astra could in September, because Codex added max after that round.
Part of BitsMinds Lab, our series of original hands-on tests.
OpenAI released GPT-6 Sol and GPT-6 Luna on 22 September at half the GPT-5.6 rates: $2 and $10 per million tokens for Sol, $0.10 and $0.50 for Luna, against $10 and $50 for GPT-6 Astra. The pitch is Astra’s gains at a fraction of Astra’s price, and on OpenAI’s own AutomationBench figures Sol beats Claude Opus 5 for 9% of the cost per task. This series asks a narrower question. It hands every model the same build briefs, allows one attempt each, and publishes whatever comes back without touching it.
The opponent is Astra itself, or rather the three files Astra produced in early September for its head-to-head with Claude Fable 5.1, the same files that later faced Grok 4.7 and Claude Opus 5.5. The briefs: a motorway interchange seen from above and a seaside fairground at dusk, each as one animated SVG, and a cinematic rocket launch as one web page. Sol and Luna ran in OpenAI’s Codex at max reasoning effort, with no browser, no preview and no way to run code. Neither could look at what it had made.
Astra won all three rounds and finished with 10 points. Sol took 5 and Luna took 3: the same order as the price list. The two cheaper models were quick, finishing each brief in between four and a half and nine minutes against Astra’s sixteen to twenty-eight. Each of their six builds has a fault that shows on the screen.
Round 1: The interchange
The brief asks for a busy interchange as one self-contained SVG, with no text and no script. Two roads cross on different levels, ramps connect them, at least twelve vehicles follow the curves of the roads they drive on in both directions, and the layering hides the lower road’s traffic under the overpass. Two traps go unstated: whether each ramp really joins the roads at both of its ends, and whether traffic keeps to the right side after it merges.
GPT-6 Astra’s interchange is the busiest this series has seen. Forty-eight vehicles, an articulated lorry and a bus among them, run along four straight through lanes and four curving ramps, and every one throws a headlight cone onto the road in front of it. Direction arrows are painted on the asphalt, which makes the traffic-direction trap easy to check by eye, and it passes. There is street lighting down the median, road signs, two ponds and a river. The bridge is a proper structure, with a deck that casts a shadow, parapets on both sides, bearings at the four corners and lower-road lane markings that stop at the edge of the deck. Every path begins off-screen, so ramp traffic arrives along the main road before it curves away, and the occlusion holds mid-crossing: a lorry comes in from the west with its cab already under the parapet and its trailer still in the open. The fault is a painted one. The white shoulder line runs straight across the exit, so the ramp never visibly separates from the road it leaves.
GPT-6 Sol built its interchange in 4 minutes 45 seconds. Twenty-four vehicles, sedans, vans and trucks with headlamps and red tail lights, drive on the ground-level east–west highway, across the north–south overpass and along four single-lane ramps. The ground is landscaped with planted verges, clumps of trees and drainage lines. The ramps are raised: each casts a shadow, the pairs that cross are drawn one above the other, and the east–west traffic passes beneath them, so the layering the brief insists on works where the ramps cross the lower road. It is their ends that fail. At the lower road each ramp starts as a rounded stub sitting in the middle of the carriageway rather than peeling off a lane, and at the other end the ramp traffic is drawn underneath the overpass deck, so cars heading up to the bridge disappear beneath it instead of joining its lanes. The ramps begin in the middle of one road and lead nowhere on the other.
GPT-6 Luna took 8 minutes 6 seconds and drew the tidiest scene of the three, closer to a planner’s map than a photograph. The east–west road is the overpass, and around the junction sit a greenbelt, local access lanes, surface car parks, garden strips, ponds and belts of trees kept clear of the roads. Eighteen vehicles run on it: six on the lower north–south road, eight on the bridge and four on the ramps. The bridge has a shadow, concrete bearings and guardrail brackets, and the ramp ends carry painted gore markings. The four ramps are broad curves that leave the lower road and climb towards the bridge, and that is as far as they get. Each one’s upper end starts under the overpass deck, in the middle of the east–west carriageway, and the cars on it are drawn beneath the deck: a ramp car travels along the bridge unseen and appears from under its edge. Luna made the same mistake as Sol, from the other direction.
Round 1 goes to Astra 6, which takes 4 points. Sol and Luna take 1 point each, because both built ramps that start in the middle of a road and connect to nothing. Astra’s only fault is a line of paint.
Round 2: The fairground
A seaside fairground at dusk, one SVG: a Ferris wheel with at least ten gondolas turning continuously and staying level with the horizon all the way round, a carousel with at least six horses rising and falling out of phase under a spinning canopy, twinkling bulbs, lights chasing round the rim of the wheel, a low sun or moon, and the whole scene reflected in the water. The gondolas are the trap the brief states. The carousel is the one it does not, and it has decided this round more often than anything else.
GPT-6 Astra’s fairground is a pastel dusk with a lighthouse and a sailboat on the water, and its carousel horses wear jewelled saddle blankets and bridles under flowing manes. They are the best-drawn horses this brief has produced. The Ferris wheel passes the stated trap cleanly: it turns once every 48 seconds while each gondola counter-rotates through a full circle over the same 48 seconds, with a gentle sway on top, so the cabins stay level. The carousel does not turn at all. Of the build’s 37 animations, thirteen are full rotations and every one of them belongs to the Ferris wheel. Eleven more are the horses bobbing on their poles and twelve are gondola sways, which leaves one: a three-degree rock back and forth about the carousel’s centre, over eight seconds. That rock is the whole ride, and the horses bob on the spot without ever travelling anywhere.
GPT-6 Sol’s scene, built in 6 minutes 18 seconds, is a pier at sunset under a band of clouds, with a low sun, a headland in the distance, a ticket kiosk, a “SWEETS” taffy booth, lamp posts, bunting and two long festoons of bulbs that twinkle independently. Its wheel stands on an A-frame and carries twelve cabins. Rim, spokes and cabins turn together once every 48 seconds while each cabin counter-rotates and sways, so the stated trap is cleared, and two rings of chasing lights run round the rim. Its carousel does turn. A striped, sloping canopy rotates every twelve seconds, and eight horses ride an elliptical track at the same speed, each on its own pole, held upright, rising and falling at staggered offsets. What they never do is turn: every horse points the same way all the way round the track, so on the far side of the ride it travels backwards. The brief also asks for the whole fairground reflected in the water, and Sol wrote one, a flattened, upside-down copy of the scene at half opacity. It then clipped that copy with a rectangle placed in the flipped coordinates, which puts the clip above the waterline and cuts every reflected ride away. The water shows ripples and nothing else. The drawing is also plainer than Astra’s.
GPT-6 Luna took 9 minutes 6 seconds over a purple dusk with a crescent moon, stars, sagging strings of warm bulbs, a “TICKETS” hut, a “SWEET TIDE” stall and a row of snack stands along a promenade of small lamps and pennants. Its Ferris wheel is a good one: twelve level cabins on a trestle, turning once every 24 seconds, with lamps chasing round the rim, and unlike Sol’s, its water holds a reflection of the rides. The carousel is where it comes apart. Its canopy rotates about the column in the flat plane of the picture, once every 14 seconds, and the platform carrying the six horses rotates separately, about a different centre, once every 18 seconds. Seen from the side, the whole ride turns like a second Ferris wheel: the horses loop over and under one another while the tilted roof spins out of step above them.
Round 2 goes to Astra 6, which takes 3 points; Sol takes 2 and Luna takes 1.
Round 3: The rocket launch
The only open brief. It names no library, asks for no 3D, and its one instruction on style is to make it look good. It wants a countdown, ignition, liftoff and a sustained ascent, with a mission-control HUD reporting the state of the flight.
GPT-6 Astra, 28 minutes.
AURORA 06 — GPT-6 Sol, 6 minutes 36 seconds.
LUNA 06 — GPT-6 Luna, 4 minutes 29 seconds. All three start on their own countdown.
GPT-6 Astra built a polished product: a three-core vehicle rendered with care, enormous editorial type, the curve of the Earth below, a flight-sequence timeline, a zoom control and a playback-speed multiplier that nobody asked for. The ascent is modelled, not decorative. Its 38.5 km at 1,003 metres per second by T+1:39 is close to a real climb, and the g-load rising from 2.13 to 3.29 as propellant burns off is the right behaviour rather than a busy-looking number. Then the flight stops developing. The sequence runs countdown, ignition, liftoff, max-Q and ascent, with no main-engine cutoff, no stage separation, no second-stage burn and no orbit. The throttle stays at 100% and the vehicle climbs until the animation runs out.
GPT-6 Sol’s AURORA 06 follows much the same plan in a quarter of the time. A three-core rocket stands beside its service tower against dark mountains, a headline reads “The sky awaits”, and a ten-second count leads into three seconds of ignition, with flame and smoke building under the pad before the rocket leaves the tower. The HUD reports altitude, velocity and downrange distance, lights up a four-step phase track, and runs a mission clock. Once the vehicle clears the tower the camera holds it at a fixed height in the frame while the ground falls away. Like Astra’s, the flight never goes past “ascent”: the engines stay at 100%, nothing separates, and the climb carries on for as long as the page stays open. During that climb the rocket’s nose sits cut off at the top of the frame. It is a smaller, plainer version of Astra’s launch with the same missing second act.
GPT-6 Luna’s LUNA 06 is the most charming of the three to look at, and on paper it has the fullest mission: an uncrewed flight to the Moon, huge thin countdown numerals, “IGN” at ignition, and a flight-status panel that steps through powered ascent to booster separation by T+22 and orbital insertion by T+47. Almost none of it is visible. The code adds the camera offset to the rocket’s position twice, once through the ground line it measures from and again on its own, so as the camera climbs the rocket slides down the screen. By about T+11 it has dropped out of the bottom of the frame, and staging and orbit play out against an empty, darkening sky. The rocket starts to rise, appears to stall, and the camera carries on into space without it.
None of the three pulled in a library, though the brief allows any from a public CDN. Counting every model that has taken this brief since September, that makes eight models from three labs drawing by hand on a 2D canvas.
Round 3 goes to Astra 6, which takes 3 points; Sol takes 2 and Luna takes 1.
Four times the speed
| Build | Astra | Sol | Luna |
|---|---|---|---|
| Interchange | 16m 15s | 4m 45s | 8m 06s |
| Fairground | 26m 33s | 6m 18s | 9m 06s |
| Rocket launch | 28m 00s | 6m 36s | 4m 29s |
| All three | 70m 48s | 17m 39s | 21m 41s |
Sol took a quarter of Astra’s time across the three briefs, and Luna about three tenths. Luna was not always the quicker of the two new models: it spent longer thinking than Sol on the interchange and the fairground, 16,331 and 12,695 reasoning tokens against Sol’s 9,020 and 8,752. Its launch was the exception. It came out of the shortest reasoning of all six runs, 2,684 tokens, and it is the build with the camera counted twice.
The conditions were the same where it mattered: the same prompts word for word, one attempt, no browser and no code execution. Two things differ. Sol and Luna ran at max, a level Codex added after Astra’s September runs at xhigh, which was then the top of the dial. And wall-clock time depends on the server as well as the model: the six Sol and Luna runs went out in parallel on one evening, Astra’s on other days. The Codex logs for Sol and Luna show nothing but file writes; neither model searched Codex’s own memory before starting, as a Codex entrant did in the second Hill Climb round.
The verdict: Astra wins with 10 points, Sol takes 5, Luna takes 3
| Interchange | Fairground | Launch | Total | |
|---|---|---|---|---|
| GPT-6 Astra | 4 | 3 | 3 | 10 |
| GPT-6 Sol | 1 | 2 | 2 | 5 |
| GPT-6 Luna | 1 | 1 | 1 | 3 |
Points were set round by round on the finished artefact alone: 4, 1 and 1 on the interchange, where two of the three builds have ramps that connect to nothing, and 3, 2 and 1 on the fairground and the launch. Astra’s builds are its September files, scored again next to the new ones. The builds were scored at BitsMinds from a local side-by-side page of the untouched outputs, with no blind scoring and no automated grading. Our GPT-6 Sol review weighs these builds against the independent benchmark numbers.
What the discount buys
Every fault in the six new builds is one a single look would catch: a ramp that ends in the middle of the road, a reflection that is not there, horses riding backwards, a rocket that leaves the shot. Astra’s builds have faults of the same kind, a painted line, a carousel that only rocks, a flight that stops at ascent, and they are fewer and smaller, set in pictures with more in them. What Sol and Luna trade away is not the idea of each build. It is the last stretch, where the pieces are joined and checked against the brief.
OpenAI’s own figures say Sol beats Claude Opus 5 on business automation for under a tenth of the cost, and nothing here contradicts that; these briefs test something else, whether a model can finish a visual build it will never see. On that test the price list turned out to be a fair guide. What the cheaper models do offer is time. At under seven minutes a brief, Sol leaves room to look at its work and send it back, which is the one thing this format forbids and most real work allows.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.