Claude Sonnet 5.5 vs Opus 5.5, Astra, Sol: A Point Short
Claude Sonnet 5.5 costs half as much as Opus 5.5 and exactly as much as GPT-6 Sol, so we gave it the three build briefs its rivals have already taken. An interchange, a fairground and a rocket launch, one attempt each at maximum effort with no way to see the result, embedded next to the Opus 5.5, GPT-6 Astra and Sol builds so you can judge them yourself.
BitsMinds LabHow this test was run4 models · 3 briefs · one attempt per model per brief · Maximum reasoning effort for every model, at the top of each one's own dial at the time: Sonnet 5.5 and Opus 5.5 in Claude Code at MAX, Astra in Codex at xhigh, Sol in Codex at maxShow details
- Models
- Claude Opus 5.5 (Claude Code) · Claude Sonnet 5.5 (Claude Code) · GPT-6 Astra (OpenAI Codex) · GPT-6 Sol (OpenAI Codex)
- Attempts
- One attempt per model per brief. No retries, no follow-up notes, no cleanup.
- Reasoning effort
- Maximum reasoning effort for every model, at the top of each one's own dial at the time: Sonnet 5.5 and Opus 5.5 in Claude Code at MAX, Astra in Codex at xhigh, Sol in Codex at max.
- Tools allowed
- Sonnet 5.5 and Opus 5.5 ran in Claude Code, Astra and Sol in OpenAI Codex, all with the same ban on opening a browser or running code. Sonnet 5.5's one shell call, before its launch, created an output folder that already existed; nothing was executed or tested.
- Timing
- Wall-clock time from prompt to finished file: Sonnet 5.5 took 49m 05s, 39m 14s and 61m 02s on 29 September; Opus 5.5 took 58m 49s, 62m 41s and 76m 33s; Astra 16m 15s, 26m 33s and 28m 00s; Sol 4m 45s, 6m 18s and 6m 36s.
- Scoring
- Points out of 5, set round by round on the finished artefact: 5, 4, 3 and 1 on the interchange, 5, 4, 3 and 2 on the fairground and the launch. Opus 5.5's, Astra's and Sol's builds are the files they were published with, scored again alongside Sonnet 5.5's. Final: Opus 5.5 wins with 13 points, Sonnet 5.5 takes 12, Astra takes 11, Sol takes 5.
- Judging
- Scored by BitsMinds from a local side-by-side comparison page of the untouched outputs. No blind scoring, no automated grading.
- Published
- September 29, 2026
Result
| Model | Provider | Score |
|---|---|---|
| Claude Opus 5.5 | Anthropic | 13 pts |
| Claude Sonnet 5.5 | Anthropic | 12 pts |
| GPT-6 Astra | OpenAI | 11 pts |
| GPT-6 Sol | OpenAI | 5 pts |
Opus 5.5 wins with 13 points, Sonnet 5.5 takes 12, Astra takes 11, Sol takes 5. Sonnet 5.5 finished second in every round at half Opus 5.5's price and beat it on the interchange; it was quicker than Opus 5.5 on every brief and far slower than either OpenAI model.
The briefs — as described in the article; the exact prompt files were not published
Round 1: Motorway interchange (SVG)
A top-down view of a busy interchange as one self-contained SVG. Two roads crossing on different levels, connecting ramps, at least twelve vehicles following the curve of the road they are on, and correct layering so traffic on the lower road is hidden by the overpass.
Result: Astra wins with 5 points, Sonnet 5.5 takes 4, Opus 5.5 takes 3, Sol takes 1: Sonnet 5.5's cloverleaf is correct everywhere, from its traffic rules to ramps that separate from the road, in a plainer picture than Astra's.
Round 2: Seaside fairground (SVG)
A seaside fairground at dusk, one SVG. A Ferris wheel with at least ten gondolas turning continuously, a spinning carousel with at least six horses rising and falling out of phase, twinkling bulbs, a reflection in the water. Each gondola must stay level with the horizon throughout the rotation.
Result: Opus 5.5 wins with 5 points, Sonnet 5.5 takes 4, Astra takes 3, Sol takes 2: Sonnet 5.5's horses travel and face the way they go but turn round in a single frame, under a wheel with lights up every spoke.
Round 3: Rocket launch (web page)
A self-contained web page that animates a rocket launching — countdown, ignition, liftoff and ascent — with a mission-control HUD. No library named, no 3D requested, one aesthetic instruction: make it look good.
Result: Opus 5.5 wins with 5 points, Sonnet 5.5 takes 4, Astra takes 3, Sol takes 2: Sonnet 5.5 flies the whole mission and deploys its satellite, under a film-grain layer and a flat band of ground that shows above the curved horizon.
Read with care
- Single run per model per brief — a re-run could land differently.
- Scored on the finished artefact only. Wall-clock time is reported but carries no points, and it depends on server load as well as the model: the four ran on different days.
- Opus 5.5's, Astra's and Sol's builds were made on earlier dates and are scored again here rather than re-run.
- Sonnet 5.5's reasoning filled Claude Code's 128,000-token limit on a single response on every brief, twice on the interchange; the harness told it to resume where it stopped, with nothing about the build.
Part of BitsMinds Lab, our series of original hands-on tests.
Claude Sonnet 5.5 is the cheaper half of Anthropic’s new generation. It was released on 28 September at $2 per million input tokens and $10 per million output, half the price of Claude Opus 5.5 and exactly the price of OpenAI’s GPT-6 Sol. On the independent Artificial Analysis index it scores two points behind Opus 5.5, and our review found it cheap and fast up to high effort and the wrong buy above it. This series asks something the benchmarks do not: hand a model the same build briefs as its rivals, allow one attempt, and publish whatever comes back untouched.
Its opponents bring the builds they were published with. Claude Opus 5.5, the model it is priced against and the leader of this series, built them for its 22 September round. GPT-6 Astra, OpenAI’s flagship, built them in early September for its head-to-head with Claude Fable 5.1. GPT-6 Sol, OpenAI’s model at Sonnet’s price, built them for its round against Astra and GPT-6 Luna. The briefs: a motorway interchange seen from above and a seaside fairground at dusk, each as one animated SVG, and a cinematic rocket launch as one web page. One attempt each, maximum reasoning effort, and no way to open a browser or run code to see the result.
Opus 5.5 won with 13 points. Sonnet 5.5 took 12, Astra took 11 and Sol took 5. Sonnet 5.5 finished second in all three rounds, one point behind the flagship that costs twice as much, and a point ahead of the one that lists at five times its price. It was quicker than Opus 5.5 on every brief and far slower than either OpenAI model; the times are in the table below.
Round 1: The interchange
The brief: a busy interchange seen from above, as one self-contained SVG with no text and no script. Two roads cross on different levels, ramps connect them, at least twelve vehicles follow the curves of the roads they drive on in both directions, and the layering hides the lower road’s traffic under the overpass. Two traps go unstated: whether each ramp really joins the roads at both ends, and whether traffic keeps to the right side after it merges.
Claude Sonnet 5.5 built a full cloverleaf. The east–west highway runs underneath, the north–south road crosses it on a bridge, and every turn between them is made the way a real cloverleaf makes it: a car that wants to turn left carries on past the bridge, peels off to the right and swings through a 270-degree loop onto the other road. Thirty-two cars and vans use it in both directions, on the correct side for right-hand traffic, and every route starts and ends off-screen, so nothing appears or disappears in view. The loops are where the layering rule bites, because a car that starts its turn under the bridge finishes it on top. Sonnet 5.5 drew every looping car twice, once on each level, and hands over from one copy to the other halfway round the loop, well away from the deck. Stepped through a full cycle in a browser, no car is ever drawn on the wrong level and none ever touches another in its lane; the tightest gap at a merge is 135 pixels. The loops also open out of the road the way ramps should. Sonnet 5.5 laid the asphalt over the kerbs, so no edge line runs across the mouth of any ramp. Around the junction sit fields, trees, an industrial yard with a car park and a housing estate with its own side street, with cloud shadows drifting across. It is a correct interchange in every respect, and a plainer picture than Astra’s.
Claude Opus 5.5 built a carefully planned one for its round, in 58 minutes 49 seconds. Four right-turn ramps, one in each quadrant, leave from the outer lane of one road and join the outer lane of the other, traffic stays on the correct side after every merge, and every path starts and ends off-screen. The merges are timed so that no car taking a ramp comes within about 145 pixels of the traffic it joins, and every ramp car carries turn signals nobody asked for. The bridge has parapets, an expansion joint at each end and a shadow thrown across the road beneath. The fault is in the paint. The ramps are drawn underneath the carriageways, and each carriageway’s solid white edge line runs straight across the point where a ramp leaves or joins it, at all eight ramp ends, so the ramps seem to slide out from under the road instead of peeling away from it.
GPT-6 Astra’s interchange, built in 16 minutes 15 seconds, is the busiest in the series. Forty-eight vehicles, an articulated lorry and a bus among them, run along four straight lanes and four curving ramps, and every one throws a headlight cone onto the road ahead. Direction arrows are painted on the asphalt, there is street lighting down the median, road signs, two ponds and a river. The bridge is a real structure, with a shadowed deck, parapets, bearings at the four corners and lower-road markings that stop at the deck’s edge, and the occlusion holds mid-crossing: a lorry comes in with its cab already under the parapet while its trailer is still in the open. Its one fault is the same line of paint as Opus 5.5’s. The white shoulder line runs straight across the exit, so the ramp never visibly separates from the road it leaves.
GPT-6 Sol built its interchange in 4 minutes 45 seconds. Twenty-four vehicles with headlamps and tail lights drive on a ground-level highway, across the overpass and along four raised single-lane ramps, over ground landscaped with verges, trees and drainage lines. The ramps cast shadows, and the layering works where they cross the lower road. Their ends do not work. Each ramp starts as a rounded stub in the middle of the lower carriageway, and at the other end its traffic is drawn underneath the overpass deck, so cars heading up to the bridge disappear beneath it instead of joining its lanes.
Round 1 goes to Astra, which takes 5 points; Sonnet 5.5 takes 4, Opus 5.5 takes 3 and Sol takes 1. Until now only Fable 5.1’s interchange, in September, had come through this brief with nothing wrong in it. Sonnet 5.5’s is the second: the traffic logic holds everywhere and its ramps separate from the road, which Opus 5.5’s do not. Astra’s picture is so much richer that it still takes the round with its painted line.
Round 2: The fairground
A seaside fairground at dusk, one SVG: a Ferris wheel with at least ten gondolas turning continuously and staying level with the horizon all the way round, a carousel with at least six horses rising and falling out of phase under a spinning canopy, twinkling bulbs, lights chasing round the rim of the wheel, a low sun or moon, and the whole scene reflected in the water. The gondolas are the trap the brief states. The carousel is the one it does not, and it keeps deciding this round.
Claude Sonnet 5.5 set its fairground on a purple dusk with a low sun on the water, striped booths, a flag on the carousel roof and strings of bulbs twinkling across the scene, all of it reflected in rippling water below. Its Ferris wheel clears the stated trap exactly: twelve gondolas ride a wheel that turns once every 40 seconds, each counter-rotating through the same full circle so that the two cancel, with a gentle six-degree sway on top. It adds a touch of its own: besides 36 bulbs chasing round the rim, a line of lights runs up every one of the twelve spokes. The carousel is drawn in perspective, as Opus 5.5’s is. Its striped skirt and roof slide round at the same speed as the horses, so roof and riders turn together, and eight horses travel round the ride on their own poles, rise and fall out of step, pass behind the centre column and in front of it, and face the way they are moving. The difference from Opus 5.5 is in how they turn round. Opus 5.5’s horses narrow as they swing round the side and turn smoothly. Sonnet 5.5’s flip to face the other way in a single frame at each end of the ride.
Claude Opus 5.5’s fairground, built in 62 minutes 41 seconds, is a full scene: a purple-to-amber sky, a crescent moon and a low sun on the water, a striped tower, strings of bulbs and booths along the pier, all mirrored in rippling water. Twelve gondolas ride a wheel that turns once every 24 seconds, each counter-rotating from its own starting angle, with a four-degree sway at staggered offsets, and the rim with its 36 bulbs turns with the spokes. Its carousel is the one that set the standard. The six horses travel round an ellipse drawn in the same proportions as the roof’s rim, narrow as they swing round the side and come back mirrored, and pass behind the centre column and in front of it. The roof’s stripes turn at exactly the horses’ speed, so the whole ride moves as one machine.
GPT-6 Astra’s, built in 26 minutes 33 seconds, is a pastel dusk with a lighthouse and a sailboat, and its carousel horses wear jewelled saddle blankets and bridles under flowing manes, the finest horses anyone has drawn for this brief. Its Ferris wheel clears the stated trap, turning once every 48 seconds while each gondola counter-rotates over the same 48 seconds with a gentle sway. Its carousel does not turn. Of the build’s 37 animations, the only one that belongs to the ride as a whole is a three-degree rock back and forth over eight seconds, so the horses bob on the spot and never travel anywhere.
GPT-6 Sol built a sunset pier in 6 minutes 18 seconds, with a ticket kiosk, a sweets booth, lamp posts, bunting and two long festoons of twinkling bulbs. Its twelve-cabin wheel stays level as it turns, with two rings of chasing lights round the rim, and its carousel does turn: a striped canopy rotates once every twelve seconds, and eight horses ride an elliptical track at the same speed, each on its own pole, rising and falling at staggered offsets. They never turn round, so on the far side of the ride every horse travels backwards. Sol also wrote the reflection the brief asks for and then clipped it away, with a rectangle placed in the flipped coordinates of the reflected scene, so the water shows ripples and nothing else.
Round 2 goes to Opus 5.5, which takes 5 points; Sonnet 5.5 takes 4, Astra takes 3 and Sol takes 2.
Round 3: The rocket launch
The only open brief. It names no library and asks for no 3D, and its one instruction on style is to make it look good. It wants a countdown, ignition, liftoff and a sustained ascent, with a mission-control HUD reporting the state of the flight.
KESTREL-1 — Claude Sonnet 5.5, 61 minutes 2 seconds. The whole flight plays in about two minutes and twenty seconds; stay for the satellite.
HALCYON-3 — Claude Opus 5.5, 76 minutes 33 seconds. Stay for orbit.
GPT-6 Astra, 28 minutes.
AURORA 06 — GPT-6 Sol, 6 minutes 36 seconds. All four start on their own countdown.
Claude Sonnet 5.5’s KESTREL-1 is the launch in this series closest to Opus 5.5’s, and it flies the whole mission. It opens on a terminal count with a launch poll of every console, then ignition at T−3, liftoff, the pitch-over and max-Q. Main-engine cutoff comes at T+2:30, with the throttle at zero and all nine engine lights out. Stage separation follows three and a half seconds later, with the spent first stage falling away on its own, then second-stage ignition at T+2:43, fairing separation at T+3:25 and second-stage cutoff at T+8:32. Then it goes a step past the point where Opus 5.5’s flight ends: the satellite separates and unfolds its solar arrays. The final screen reports a 203 by 203 km orbit at 7.79 km/s, the right speed for a circular orbit at that height. Mission time runs up to fifteen times faster through the long burn and slows for every event, with the rate shown on screen, so the whole flight plays in about two minutes and twenty seconds. It launches at night under floodlights, the camera keeps the vehicle in frame from the pad to orbit, and the mission-control screen is the densest in the series: altitude, velocity, Mach, downrange, g-load, dynamic pressure, pitch, throttle, propellant for each stage, engine status, three live plots and a caption for every call from the consoles.
Two things spoil the picture. A film-grain layer is drawn over the whole frame, and it reads as noise. And once the flight is high enough for the Earth to curve, the ground is still painted as a flat band from the horizon line down, so a straight-edged strip of it shows above the curved horizon at the sides of the picture. It looks like a bug because it is one.
Claude Opus 5.5’s HALCYON-3, built in 76 minutes 33 seconds, flies the entire mission too. It has a terminal count with a go/no-go poll of every station, ignition at T−3, liftoff, the pitch program, a throttle-down through max-Q, main-engine cutoff at T+2:00, stage separation, second-stage ignition at T+2:10, the spent booster’s flip, fairing separation and the Kármán line. At T+7:53 the second stage shuts down into a 203 by 207 km orbit, holding 28,024 km/h at 205 km, within a few km/h of circular speed at that height. It is shot like a film: it launches in twilight, cuts between camera angles, pushes in on the engines at ignition as the exhaust rolls across the pad, and pulls back as the vehicle climbs, and the camera never loses the vehicle. The flight plays in about two minutes.
GPT-6 Astra built a polished product in 28 minutes, and a flight that stops early: a three-core vehicle rendered with care, enormous editorial typography, the curvature of the Earth below, a flight-sequence timeline, a zoom control and a playback-speed multiplier nobody asked for. Its ascent is modelled rather than decorative. It reaches 38.5 km at 1,003 metres per second by T+1:39, and its g-load climbs from 2.13 to 3.29 as the propellant burns away. And then it stops developing. Its sequence reads countdown, ignition, liftoff, max-Q and ascent, with no main-engine cutoff, no stage separation, no second-stage burn and no orbit. The throttle stays at 100% until the animation runs out.
GPT-6 Sol’s AURORA 06, built in 6 minutes 36 seconds, follows much the same plan in a quarter of the time. A three-core rocket stands beside its service tower against dark mountains, a ten-second count leads into three seconds of ignition, and the HUD reports altitude, velocity and downrange distance on a four-step phase track. Once the vehicle clears the tower, the camera holds it at a fixed height while the ground falls away. Like Astra’s, the flight never goes past “ascent”: the engines stay at 100%, nothing separates, and the rocket’s nose sits cut off at the top of the frame for the whole climb.
Round 3 goes to Opus 5.5, which takes 5 points; Sonnet 5.5 takes 4, Astra takes 3 and Sol takes 2.
None of the four pulled in a library, though the brief allows any from a public CDN. Counting every model that has taken this brief since September, that makes ten models from three labs drawing their launches by hand on a 2D canvas.
Quicker than Opus, a long way behind OpenAI
| Build | Sonnet 5.5 | Opus 5.5 | Astra | Sol |
|---|---|---|---|---|
| Interchange | 49m 05s | 58m 49s | 16m 15s | 4m 45s |
| Fairground | 39m 14s | 62m 41s | 26m 33s | 6m 18s |
| Rocket launch | 61m 02s | 76m 33s | 28m 00s | 6m 36s |
| All three | 149m 21s | 198m 03s | 70m 48s | 17m 39s |
Like Opus 5.5, Sonnet 5.5 spent almost all of that time thinking. On every brief its first response was 128,000 tokens of reasoning and nothing else, the most Claude Code allows in a single response, and on the interchange it filled that limit twice. It started writing the interchange after 48 minutes and finished the file in under one. It started the fairground after 32 minutes and spent seven and a half writing and correcting it. The launch took 36 minutes of thinking, then 25 more in which it wrote the page in pieces and read its own code back eleven times, the only kind of checking these rules allow. It is the same verbosity at maximum effort that our review found in the independent measurements.
The conditions matched where it mattered. Sonnet 5.5 received Opus 5.5’s prompts word for word apart from the save path, and like Opus 5.5 it ran in Claude Code at maximum effort with no browser and no way to run code. Astra ran in OpenAI’s Codex at the highest setting Codex then offered, and Sol at max. There was one small departure: before writing its launch, Sonnet 5.5 made a single shell call that created the output folder, which already existed. It ran nothing else, and none of its three builds was ever executed or tested.
The verdict: Opus 5.5 wins with 13 points, Sonnet 5.5 takes 12
| Interchange | Fairground | Launch | Total | |
|---|---|---|---|---|
| Claude Opus 5.5 | 3 | 5 | 5 | 13 |
| Claude Sonnet 5.5 | 4 | 4 | 4 | 12 |
| GPT-6 Astra | 5 | 3 | 3 | 11 |
| GPT-6 Sol | 1 | 2 | 2 | 5 |
Points were set round by round on the finished artefact alone, out of five: 5, 4, 3 and 1 on the interchange, and 5, 4, 3 and 2 on the fairground and the launch. Opus 5.5’s, Astra’s and Sol’s builds are the files they were published with, scored again alongside Sonnet 5.5’s. The builds were scored at BitsMinds from a local side-by-side page of the untouched outputs, with no blind scoring and no automated grading.
Half the price, one point back
Sonnet 5.5 costs half as much as Opus 5.5 per token and finished a single point behind it. It finished a point ahead of Astra, which lists at five times its price, and seven ahead of Sol, which costs the same. It lost to Opus 5.5 on finish rather than on understanding. Its horses turn round in one frame instead of swinging, and its launch carries grain and a stray band at the horizon. Its interchange, meanwhile, is the second in the series with nothing wrong in it, where Opus 5.5’s has a line across every ramp end. What it did not save was time: it thought for 32 to 48 minutes on every brief before writing a line, and filled Claude Code’s reasoning limit each time.
For a Sonnet, it is a new position. Its predecessor finished last in both of the Claude rounds it entered, in August and in September, several points behind the Opus it ran against. Sonnet 5.5 beat the flagship on the interchange and finished a point behind it on the other two, and it is the only model in this field that never finished lower than second.
More on Claude
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.