Claude Haiku 5.5 vs GPT-6 Luna: Lost in Thought
Claude Haiku 5.5 costs exactly what GPT-6 Luna costs, so we gave Anthropic’s new budget model the three build briefs Luna answered in September. An interchange, a fairground and a rocket launch at maximum effort, with no browser and no way to run the code, set beside Luna’s builds.
BitsMinds LabHow this test was run2 models · 3 briefs · each attempt was one shot: no follow-up notes, no corrections, no cleanup · Luna in OpenAI Codex at max on all three briefs. Haiku 5.5 in Claude Code at max, where it finished only the fairground; its interchange comes from a run at xhigh, one step down, after two attempts at max produced no file, and its launch produced no file at max or at xhighShow details
- Models
- GPT-6 Luna (OpenAI Codex) · Claude Haiku 5.5 (Claude Code)
- Attempts
- Each attempt was one shot: no follow-up notes, no corrections, no cleanup. Haiku 5.5 had two attempts at maximum effort on the interchange and the launch, and neither produced a file; it then ran both at xhigh, where it finished the interchange, which is scored, and again produced no launch.
- Reasoning effort
- Luna in OpenAI Codex at max on all three briefs. Haiku 5.5 in Claude Code at max, where it finished only the fairground; its interchange comes from a run at xhigh, one step down, after two attempts at max produced no file, and its launch produced no file at max or at xhigh.
- Tools allowed
- Haiku 5.5 ran in Claude Code and Luna in OpenAI Codex, both with the same ban on opening a browser or running code. Haiku 5.5's transcripts show nothing but writes and edits to its own output file, apart from one search for that file's name before it wrote the interchange.
- Timing
- Wall-clock time from prompt to finished file: Haiku 5.5's fairground took 16m 32s at max and its interchange 12m 32s at xhigh, on 8 October. Its first maximum-effort attempts on the interchange and the launch reasoned for 37m 31s and 39m 40s without producing a file, and the launch ran 36m 40s at xhigh without one. Luna took 8m 06s, 9m 06s and 4m 29s on 22 September.
- Scoring
- Points set round by round on the finished artefact. The interchange and the fairground were scored head to head: on the interchange Luna took 2 points and Haiku 5.5 took 1, for the build it made at xhigh; on the fairground Haiku 5.5 took 2 and Luna took 1. On the launch Luna keeps the 1 point its build took on 23 September, and Haiku 5.5, with no file, takes 0. Final: Luna wins with 4 points, Haiku 5.5 takes 3.
- Judging
- Scored by BitsMinds from a local side-by-side comparison page of the untouched outputs. No blind scoring, no automated grading.
- Published
- October 8, 2026
Result
| Model | Provider | Score |
|---|---|---|
| GPT-6 Luna | OpenAI | 4 pts |
| Claude Haiku 5.5 | Anthropic | 3 pts |
Luna wins with 4 points, Haiku 5.5 takes 3. Haiku 5.5 won the fairground, the one build it finished at maximum effort. On the interchange and the launch it filled Claude Code's 128,000-token response limit with reasoning again and again; one step down it finished the interchange, and it never produced a launch.
The briefs — as described in the article; the exact prompt files were not published
Round 1: Motorway interchange (SVG)
A top-down view of a busy interchange as one self-contained SVG. Two roads crossing on different levels, connecting ramps, at least twelve vehicles following the curve of the road they are on, and correct layering so traffic on the lower road is hidden by the overpass.
Result: Luna wins with 2 points, Haiku 5.5 takes 1: at maximum effort Haiku 5.5 handed in no file, and its xhigh build, the plainest picture of the series, has ramps that leave the lower road at right angles and loop back to it under the bridge, with cars fading in and out; Luna's ramps climb towards the bridge and start under its deck.
Round 2: Seaside fairground (SVG)
A seaside fairground at dusk, one SVG. A Ferris wheel with at least ten gondolas turning continuously, a spinning carousel with at least six horses rising and falling out of phase, twinkling bulbs, a reflection in the water. Each gondola must stay level with the horizon throughout the rotation.
Result: Haiku 5.5 wins with 2 points, Luna takes 1: both wheels keep their cabins level; Haiku 5.5's horses rise and fall in a row without going round, while Luna's carousel spins like a second Ferris wheel and its lowest cabins sink into the boardwalk.
Round 3: Rocket launch (web page)
A self-contained web page that animates a rocket launching — countdown, ignition, liftoff and ascent — with a mission-control HUD. No library named, no 3D requested, one aesthetic instruction: make it look good.
Result: Luna wins with 1 point, Haiku 5.5 takes 0: Haiku 5.5 handed in no file at maximum effort or at xhigh; Luna plans staging and orbit, but its camera offset is counted twice and the rocket leaves the frame by T+11.
Read with care
- Luna's builds are the files it was published with on 23 September, scored again here rather than re-run; on the launch its point is the one it took then.
- Haiku 5.5's interchange was built at xhigh, one step below its top setting, after two attempts at maximum effort produced no file; its launch produced no file at max or at xhigh and scores 0.
- Scored on the finished artefact only. Wall-clock time is reported but carries no points, and it depends on server load as well as the model: the two ran on different days.
- Haiku 5.5 ran in Claude Code and Luna in OpenAI Codex, so their limits differ: Claude Code caps a single response at 128,000 output tokens.
Part of BitsMinds Lab, our series of original hands-on tests.
Claude Haiku 5.5 is the smallest model in Anthropic’s new generation and, since 7 October, the cheapest: $0.10 per million input tokens and $0.50 per million output for prompts up to 100,000 tokens. That is a tenth of Haiku 4.5’s price and exactly what OpenAI charges for GPT-6 Luna. On the Artificial Analysis Intelligence Index it scores 43 at maximum effort, five points above Luna’s 38, though each index task costs $0.21 against Luna’s $0.07. It is also the first Haiku with an effort setting, from Low to Max, and our review follows it across all five levels.
Its opponent brings the builds it was published with. GPT-6 Luna answered the same three briefs on 22 September for OpenAI’s family round, where it finished behind GPT-6 Astra and GPT-6 Sol: a motorway interchange and a seaside fairground, each as one animated SVG, and a cinematic rocket launch as one web page. Haiku 5.5 received Luna’s prompts word for word apart from the save path and ran in Claude Code at maximum effort, with no browser, no way to run code and nothing sent after the prompt.
Luna won with 4 points and Haiku 5.5 took 3. Haiku 5.5 took the fairground, Luna took the interchange, and the launch went to Luna unopposed. At maximum effort Haiku 5.5 finished only the fairground: on the other two briefs it reasoned until it reached the 128,000-token limit Claude Code puts on a single response, again and again, and never started writing. One step down the dial it finished the interchange, and that build is scored here. It never finished the launch at all.
Round 1: The interchange
The brief asks for a busy interchange as one self-contained SVG, with no text and no script. Two roads cross on different levels, ramps connect them, at least twelve vehicles follow the curves of the roads they drive on in both directions, and the layering hides the lower road’s traffic under the overpass. Two traps go unstated: whether each ramp really joins the roads at both of its ends, and whether traffic keeps to the right side after it merges.
Claude Haiku 5.5 handed in nothing at maximum effort. Its first attempt reasoned for 37 minutes 31 seconds over four responses in a row, each one stopping at the 128,000-token limit with no code in it, after which Claude Code ended the run. A second attempt, started from scratch, opened with another 128,000 tokens of reasoning and no code. One step down the dial, at xhigh, it finished in 12 minutes 32 seconds, and that is the build scored here.
It is the plainest picture any model has drawn for this brief: flat green ground, eighteen round trees, two grey roads and a shadow under the bridge, with no signs, no parapets and no lane arrows. Eighteen vehicles keep to the right on both roads, and the lower road’s traffic passes correctly under the bridge. The ramps are where it goes wrong. All four leave the lower road at right angles, bend round under the bridge deck and come back down to the same lower road, so they form two loops at ground level that never reach the upper road. The cars on them fade into view and out again, two of them in the middle of the open highway.
GPT-6 Luna took 8 minutes 6 seconds and drew a scene closer to a planner’s map than a photograph. The east–west road is the overpass, and around the junction sit a greenbelt, local access lanes, surface car parks, garden strips, ponds and belts of trees kept clear of the roads. Eighteen vehicles use it: six on the lower north–south road, eight on the bridge and four on the ramps. The bridge has a shadow, concrete bearings and guardrail brackets, and the ramp ends carry painted gore markings. The four ramps are broad curves that leave the lower road and climb towards the bridge, and that is as far as they go. Each one’s upper end begins under the overpass deck, in the middle of the east–west carriageway, and the cars on it are drawn beneath the deck, so a ramp car travels along the bridge unseen and appears from under its edge.
Round 1 goes to Luna, which takes 2 points; Haiku 5.5 takes 1. Neither has ramps that join the two roads, and Haiku 5.5’s are bent at right angles, carry cars that appear and vanish in the open and sit in the barest picture of the two.
Round 2: The fairground
A seaside fairground at dusk, one SVG: a Ferris wheel with at least ten gondolas turning continuously and staying level with the horizon all the way round, a carousel with at least six horses rising and falling out of phase under a spinning canopy, twinkling bulbs, lights chasing round the rim of the wheel, a low sun or moon, and the whole scene reflected in the water. The gondolas are the trap the brief states. The carousel is the one it does not: canopy, platform and horses should turn round the centre pole together, the horses facing the way they travel.
Claude Haiku 5.5 took 16 minutes 32 seconds over a sunset pier: a sky running from deep navy to orange, a low sun sitting on the horizon with its glitter on the water, a plank pier on posts, two strings of bulbs slung from one side of the picture to the other, and the whole fairground mirrored in the sea below. Its Ferris wheel clears the stated trap exactly. Twelve cabins hang from a wheel that turns once every 24 seconds, each counter-rotating through the same full circle so that it stays level, with a three-degree sway on top, and the lights on the rim turn with the wheel. The carousel is where it stops short. Its red-and-white canopy seems to turn because the stripes slide sideways inside the dome, but nothing underneath goes round: six white horses stand in a row on fixed poles, all facing right, rising and falling on the spot half a second out of step with one another.
GPT-6 Luna took 9 minutes 6 seconds over a purple dusk with a crescent moon, stars, sagging strings of warm bulbs, a “TICKETS” hut, a “SWEET TIDE” stall and a row of snack stands along a promenade of small lamps and pennants. Its wheel clears the stated trap too, with twelve level cabins on a trestle, a turn every 24 seconds and lamps chasing round the rim, and its water holds a reflection of the rides. Its carousel comes apart. The canopy rotates about the column in the flat plane of the picture, once every 14 seconds, and the platform carrying the six horses rotates separately, about a different centre, once every 18 seconds, so the whole ride turns like a second Ferris wheel, its horses looping over and under one another beneath a tilted roof that spins out of step. Set beside Haiku 5.5’s, two more things show. Luna’s wheel hangs lower than the boardwalk it stands on, so at the bottom of every turn the lowest cabin sinks through the deck towards the water, and its strings of bulbs end in mid-air, tied to nothing.
Round 2 goes to Haiku 5.5, which takes 2 points; Luna takes 1. Neither carousel goes round its centre pole the way the best builds of this brief do, but Haiku 5.5’s stays upright under its roof, its cabins stay out of the water and its bulbs do not float.
Round 3: The rocket launch
The only open brief. It names no library and asks for no 3D, and its one instruction on style is to make it look good. It wants a countdown, ignition, liftoff and a sustained ascent, with a mission-control HUD reporting the state of the flight.
LUNA 06 — GPT-6 Luna, 4 minutes 29 seconds. Keep watching after T+11.
Claude Haiku 5.5 never got as far as a countdown. At maximum effort its first attempt reasoned for 39 minutes 40 seconds over four responses, each stopping at the 128,000-token limit with no code in it, before Claude Code ended the run, and a second attempt opened the same way. At xhigh it did it again: four responses at the limit, 36 minutes 40 seconds, no file. Nine responses across the three attempts, more than a million tokens of reasoning, and not a line of HTML.
GPT-6 Luna’s LUNA 06 is charming to look at, and on paper it plans the fullest mission of the cheap models in this series: an uncrewed flight to the Moon, huge thin countdown numerals, “IGN” at ignition, and a flight-status panel that steps through powered ascent to booster separation by T+22 and orbital insertion by T+47. Very little of it is visible. The code adds the camera offset to the rocket’s position twice, once through the ground line it measures from and again on its own, so as the camera climbs the rocket slides down the screen. By about T+11 it has dropped out of the bottom of the frame, and the staging and the orbit play out against an empty, darkening sky. It came out of the shortest reasoning of any of Luna’s runs, 2,684 tokens.
Luna pulled in no library, though the brief allows any from a public CDN, and neither has any model that has finished this brief since September. Haiku 5.5 never got far enough to choose.
Round 3 goes to Luna, with the 1 point its build took in September; Haiku 5.5 takes 0.
Half a million tokens, and no file
| Build | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Interchange | 12m 32s at xhigh; no file at max | 8m 06s |
| Fairground | 16m 32s | 9m 06s |
| Rocket launch | no file at max or at xhigh | 4m 29s |
| All three | — | 21m 41s |
| API list price, per million tokens in / out | $0.10 / $0.50 | $0.10 / $0.50 |
| Intelligence Index (v4.3.2, max) | 43 | 38 |
| Cost per index task | $0.21 | $0.07 |
Claude Code caps a single response at 128,000 output tokens. When a model fills it, Claude Code tells it to carry on where it stopped, and after the fourth stop in a row it ends the run. The other Claude models in this series have filled that cap on these briefs too: Opus 5.5 once on each of them, and Sonnet 5.5 once on two and twice on the interchange, and both then wrote their files. Haiku 5.5 filled it once on the fairground and then wrote the whole scene in about seven minutes. On the interchange and the launch, at maximum effort, it never got past reasoning. At xhigh it filled the limit once on the interchange before writing it, and four times again on the launch.
It is not a question of money. At Haiku 5.5’s list price, the 512,000 output tokens of a failed attempt cost about 26 cents. It is a question of time, and of where on the dial to run the model: Artificial Analysis puts Haiku 5.5 at High level with Luna’s best score, 38, for $0.08 a task. Whether High would have finished these briefs, this test does not show. Luna, at the top of its own dial, reasoned for between 2,684 and 16,331 tokens per brief and wrote every file in under ten minutes.
The conditions were the same where it mattered: Luna’s prompts word for word apart from the save path, no browser and no code execution for either model, and nothing sent after the prompt. Luna ran at max on all three briefs; Haiku 5.5’s fairground comes from max and its interchange from xhigh. The two ran in different harnesses, Haiku 5.5 in Claude Code and Luna in OpenAI’s Codex, and the 128,000-token limit on a single response is Claude Code’s. Wall-clock time depends on the server as well as the model, and the two ran on different days. Haiku 5.5’s transcripts show nothing but writes and edits to its own output file, apart from one search for that file’s name before it wrote the interchange.
The verdict: Luna wins with 4 points, Haiku 5.5 takes 3
| Interchange | Fairground | Launch | Total | |
|---|---|---|---|---|
| GPT-6 Luna | 2 | 1 | 1 | 4 |
| Claude Haiku 5.5 | 1 | 2 | 0 | 3 |
Points were set round by round on the finished artefact alone: 2 and 1 on the interchange and on the fairground, while on the launch Luna keeps the point its build took in September and Haiku 5.5, with no file at max or at xhigh, takes none. Haiku 5.5’s interchange is the one it built at xhigh, after two attempts at maximum effort produced nothing. Luna’s builds are the files it was published with, scored again beside Haiku 5.5’s, which is why its interchange now takes 2 points where it took 1 in September. The builds were scored at BitsMinds from a local side-by-side page of the untouched outputs, with no blind scoring and no automated grading. Our Claude Haiku 5.5 review has the benchmark numbers at all five effort settings.
Same price, one file short
Haiku 5.5’s fairground beats Luna’s, and it clears the stated trap with the same arithmetic the bigger models use. Its interchange, built a step below its top setting, is the plainest of the series and repeats the mistake Luna and GPT-6 Sol made in September. Its launch does not exist.
What separates the two models is not what they can draw but whether they stop thinking long enough to draw it. Luna’s three files are finished and each is wrong somewhere you can see; at the top of its dial, Haiku 5.5 finished one. At this price either model is cheap enough to run twice, but in this test only Luna handed in something to look at every time.
More on Claude
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.