Claude Opus 5.5 vs GPT-6.1 Sol: A Familiar Hill Climb
GPT-6.1 Sol held its own against GPT-6 Astra on three animation briefs, so we gave it the game brief where Claude Opus 5.5 set the bar. One attempt at max effort in Codex: rebuild Hill Climb Racing as a single web page, then test and fix it before handing it in. Both games are here for you to drive.
BitsMinds LabHow this test was run2 models · 1 brief · one attempt per model: each was required to run, exercise and fix its own build before reporting done, and the report froze the file · Claude Opus 5.5 at maximum in Claude Code; GPT-6.1 Sol at max in OpenAI Codex, not ultra, which adds automatic task delegationShow details
- Models
- Claude Opus 5.5 (Claude Code) · GPT-6.1 Sol (OpenAI Codex)
- Attempts
- One attempt per model: each was required to run, exercise and fix its own build before reporting done, and the report froze the file. GPT-6.1 Sol's prompt was checked to differ from the one Opus 5.5 received on exactly two lines, both of them output paths.
- Reasoning effort
- Claude Opus 5.5 at maximum in Claude Code; GPT-6.1 Sol at max in OpenAI Codex, not ultra, which adds automatic task delegation.
- Tools allowed
- Both were told to verify before reporting done, and both had a shell, Node and a browser to do it with. Each ran inside its own maker's agentic coding environment.
- Timing
- Wall-clock time from prompt to finished file: GPT-6.1 Sol took 57 minutes 38 seconds on 30 September; Claude Opus 5.5 took 98 minutes 6 seconds on 23 September.
- Scoring
- Four categories scored out of five: graphics, physics, sound and gameplay. Opus 5.5's build is its 23 September file with its scores unchanged. Final: Claude Opus 5.5 wins with 19 points, GPT-6.1 Sol takes 13. No category for faithfulness to the original game and no credit for the amount of testing a model did.
- Judging
- Scored by BitsMinds from a local side-by-side comparison page of the untouched outputs. No blind scoring, no automated grading. Both games were played in an ordinary browser, with GPT-6 Astra's Hill Climb build open beside them for reference. The measurements quoted in the article — crash distances, coin counts and fuel at fixed speeds, suspension travel, jump lengths — were taken afterwards by driving each delivered build's own code and reading its own state.
- Published
- September 30, 2026
Result
| Model | Provider | Score |
|---|---|---|
| Claude Opus 5.5 | Anthropic | 19 pts |
| GPT-6.1 Sol | OpenAI | 13 pts |
Claude Opus 5.5 wins with 19 points, GPT-6.1 Sol takes 13. GPT-6.1 Sol built a far longer and harder game than GPT-6 Sol did, but one whose coins and fuel ask little of the player, and whose layout and physics constants match GPT-6 Astra's earlier build.
The briefs — as described in the article; the exact prompt files were not published
Round 1: Hill Climb Racing (single HTML file)
The brief first run on 19 September, word for word apart from the output path: recreate Hill Climb Racing as one self-contained HTML file — everything inline, no CDN, no network, no images, all art drawn in code — surviving an iframe sandboxed with allow-scripts only. Three screens in a loop with no dead ends, a two-wheeled vehicle with suspension over hilly terrain, touch and keyboard controls, a real fail state, a restart without a reload, and physics not tied to the frame rate. Fuel and coins are deliberately not mentioned.
Result: Opus 5.5 wins with 19 points, GPT-6.1 Sol takes 13. GPT-6.1 Sol's 900-metre course rolls a buggy driven flat out, but every finishing run collects all 80 coins and its fuel only runs out on a car that has stalled; its game is a close relative of the one GPT-6 Astra built for the first field.
Read with care
- Single run per model — a re-run could land differently.
- Two different harnesses (Claude Code and Codex) mean the environment, not just the model, was under test.
- Opus 5.5's build was made on 23 September and is carried over rather than re-run.
- Neither entrant started cold: GPT-6.1 Sol read its Codex memory of earlier one-shot builds before writing code, and Claude Code gave the Opus 5.5 subagent this project's memory index. Neither contains anything about this brief.
Part of BitsMinds Lab, our series of original hands-on tests.
The Hill Climb brief is the one game in this series: rebuild Hill Climb Racing from a single prompt, as one self-contained HTML file a reader can play. Seven models from three labs had taken it before this round. The bar was set on 23 September, when Claude Opus 5.5 built the first entry that is fun to play for its own sake, in its round against GPT-6 Sol and Grok 4.7.
OpenAI’s GPT-6.1 Sol arrived on 29 September with the claim that it nearly matches GPT-6 Astra at a fifth of the price, and on our three animation briefs it finished level with Astra. So it gets the game brief as well, against the build that set the bar. It received the prompt Opus 5.5 was given, word for word apart from the output path, ran once in OpenAI’s Codex at max reasoning effort, and was told to run, test and fix its game before handing it in. Opus 5.5 brings the file it built on 23 September, unchanged.
Opus 5.5 wins with 19 points out of 20, and GPT-6.1 Sol takes 13. The bigger surprise was the game GPT-6.1 Sol built. It is a close relative of the one GPT-6 Astra made for the first field on 19 September: the same screen layout, the same scale of level and, inside the code, the same physics constants.
Play them both
Both games run below exactly as the models delivered them. On a desktop, click into a game first so it receives the keys, then drive with the arrow keys or WASD. On a phone, use the on-screen pedals. GPT-6.1 Sol’s game starts with its sound switched off; the speaker button turns it on.
Hill Hopper — Claude Opus 5.5, 98 minutes 6 seconds.
Ridge Run — GPT-6.1 Sol, 57 minutes 38 seconds.
The brief
One HTML file with everything inline: no CDN, no network requests, no images, no web fonts, and every piece of art drawn in code. It has to work inside an iframe sandboxed with allow-scripts only, with no local storage and no cookies, because that is how it sits on this page. It needs three screens, an opening screen, a menu and the game with one level, linked in a loop with no dead ends. In the game: a two-wheeled vehicle on suspension over hilly terrain, a camera that follows it, throttle and brake on the keyboard and as touch controls, wheels that stay on the ground, a car that can leave it over a crest and flip on a bad landing, a fail state, a restart without a reload, a HUD, an end-of-run screen, and physics that does not depend on the frame rate. Last of all, it asks for a game someone would actually want to play.
Fuel and coins are never mentioned. They are the two systems the original game is built on, every model that has taken the brief has shipped collectibles of some kind unprompted, and whether they ask anything of the player has decided every field so far.
Hill Hopper, by Claude Opus 5.5
Opus 5.5 took 98 minutes 6 seconds, longer than any other model has spent on this brief, and delivered the largest file: 113 KB over 2,072 lines. Its title screen has the buggy driving itself behind an animated “HILL HOPPER” logo. The menu shows the level, Pinecrest Ridge, with an elevation profile, the session’s best distance, best time and banked coins, five paint jobs to choose from, a help card and a sound switch. The world is a bright cartoon day: a snow-capped range, two wooded hill layers, dirt with rock strata under a grass edge, and a red buggy with a roll hoop, coil-over struts and a helmeted driver in a scarf. The HUD carries a fuel gauge, a coin counter, a progress bar with the fuel cans marked on it, a timer and a speedometer.

The course runs 1,176 metres in nine named sections, from The Meadows to the Final Climb, with a washboard, a big jump over a valley and a row of camel humps, and it asks something of the driver everywhere. Flat out, the buggy reaches 80 km/h and lands on its driver’s helmet at about 260 metres, which ends the run. Five fuel cans are spread along the way and a full tank lasts 25 seconds at full throttle, so a driver who crawls along at about 12 km/h runs dry at 190 metres, just short of the first can. Of its 214 coins, 60 hang in ten arcs over the humps and kickers, up to about five metres off the ground, along the path the buggy flies at the right speed. Held at around 29 km/h, a run collects 18 of those 60. Held at around 50 km/h, with no steering in the air at all, it collects 56.
The driving is what makes it. The buggy is quick and light, and in the air the pedals tip it back and forward, gently for a tap and far enough to flip for a long press. Flips and long jumps pay bonus coins. The suspension does real work: pulling away from rest, the rear spring compresses 4.6 centimetres while the front extends 4.9 and the nose lifts, and braking from 32 km/h reverses it, the front compressing 6.2 centimetres and the rear extending 7.1, before both settle back exactly once the car has stopped. The tyres throw dust and clods of earth in proportion to how much they slip, with up to 23 dust clouds and 14 clods in the air at once as the buggy pulls away. Lift off the throttle and it coasts, from 73 km/h to 51 km/h over three seconds, while the camera stays ahead of the car and pulls back as the speed rises.
Opus 5.5 also tested the course, not only the engine. The scripts it wrote include a driver crawling at six different speeds to find where the tank runs out, a scan of the whole course for every slope steeper than 38 degrees with six driving styles sent through each one, and a full run played on real key presses. They found six bugs, all listed as fixed in its report, among them “fuel never ran out, even at a crawl” and a 50-degree drop that crashed every driving style starting from rest. It left one thing in the file that does not belong there: a window.__HH test hook that includes an autopilot switch. The report discloses it, and it has no effect on normal play.
Ridge Run, by GPT-6.1 Sol
GPT-6.1 Sol took 57 minutes 38 seconds, almost twice as long as GPT-6 Sol spent on this brief, and delivered a 58 KB file of 194 lines. It opens on an editorial title page: “RIDGE RUN.” in heavy italic capitals over an outlined second line, an invented “Summit Motor Co.” badge, the line “A little engine. A big mountain.”, and an orange buggy parked on a crest in front of snow-capped peaks. The menu is a single card for the level, Pinecrest Pass, with an elevation profile, its length and climb, 900 metres and 76 metres, and a key guide that explains the pedals double as air control. The world is a soft sage-and-teal mountain range with pines, drifting clouds and a pale sun, and the buggy is an orange roll-cage runabout with a driver in a green helmet. The HUD has cards for distance, coins and fuel with the name of the current section, a progress bar counting down the metres to go, pause and sound buttons, a speedometer and a timer. The end card reads “Hello, summit.”, “That went sideways.” or “Running on empty.”, and reports distance, time, coins and air time, with the session’s best underneath.

The course is 900 metres long in five named sections, from The Meadows to the Summit Approach, and it has to be driven with some care. Held flat out, the buggy reaches 65 km/h and rolls onto its cage at 378 metres, 29 seconds in. Held to 60 km/h, it gets further and still rolls, at 594 metres. Held at 30, 40 or 50 km/h, it reaches the summit, in 81 to 115 seconds. The jumps are short: a run held at 50 km/h leaves the ground five times for longer than three tenths of a second, and the longest lasts just over a second.
Coins and fuel are both there, unprompted, and neither asks much of the player. The 80 coins come in twenty groups of four, each hovering 1.7 to 2 metres above the ground beneath it, well inside the 2.1-metre radius at which the buggy picks them up. Every run that reaches the summit collects all 80, including one held at 30 km/h that never leaves the ground. Five fuel cans each add 42%, and a full tank lasts 80 seconds at full throttle, most of a run to the summit. On every finishing run we measured, the gauge never fell below 76%. The fuel does end runs, but only runs that have already stopped: held at 8, 12 or 20 km/h, the buggy stalls on a steep climb, at 252 or 467 metres, and sits there with the throttle held until the tank is empty.
The physics underneath is sound and has some weight to it. It runs at a fixed 120 Hz step, and the suspension moves the way it should: pulling away, the rear spring compresses 10 centimetres while the front extends by as much, and braking from 32 km/h stops the buggy in just over a second, with the front compressing 12 centimetres and the rear extending 13.7. In the air, the pedals tip the nose up and down. The camera pulls back by up to 13% as the speed rises, the rear tyre kicks up dust at speed and landings raise a puff of dirt. Once switched on, the sound is an engine tone that rises with speed, a chime for each coin, a tone for each fuel can and a closing note, high for the summit and low for a failure.
GPT-6.1 Sol tested a great deal. It drove the game in headless Chrome through full runs, a rollover and an empty tank, restarts and multi-touch pedals, took about two dozen screenshots on desktop and phone layouts, ran the game inside a sandboxed iframe with storage blocked, stress-tested the physics and checked that the simulation gives the same result at different frame rates. Its own tests included the one that matters most for this level: held flat out, its buggy rolls at 378 metres, and a driver who feathers the throttle finishes. The report closed with “No known unfinished or broken features.” It also left its test folder, scripts and screenshots, next to the file it delivered. Only the game itself is embedded on this page.
A game we had seen before
Played next to GPT-6 Astra’s Ridgeline, from the first field, Ridge Run looks like a second draft of the same game. The HUD is laid out the same way: distance, fuel and coin cards along the top left, pause and sound buttons top right, brake and gas as two cards in the bottom corners, a speed readout in km/h in the middle of the bottom edge, and a message across the centre of the screen when a fuel can is collected. Both are drawn in flat editorial colour with a sun disc, a few birds and a small buggy with a roll cage. Both have exactly five fuel stops, and both courses are just under a kilometre long: Astra’s runs to 960 metres and Sol’s to 900.

The code is closer still. Both files open their simulation with a single line declaring a fixed step of 1/120 of a second, then the finish distance, then the start position: const DT=1/120, FINISH=960, START=8 in Astra’s, const DT=1/120, FINISH=907, START=7 in Sol’s. Both give the buggy’s body a mass of 210 and its wheels a radius of 0.43 metres. No other build this brief has produced has that pair. GPT-5.6 Sol, another OpenAI model, also chose 0.43-metre wheels, under a body with a mass of 6.4. The builds from Claude Opus 5, Opus 5.5, Fable 5.1 and Grok 4.7 share neither figure: Opus 5.5, for one, steps its physics at 240 Hz with a body mass of 170.
Nothing in GPT-6.1 Sol’s Codex log suggests it saw Astra’s file. None of the commands it ran opens Astra’s build or any earlier entry, and it made no web searches. What it did read, before writing any code, was Codex’s own memory of earlier one-shot builds, which holds nothing about this brief. The likeliest explanation is lineage: the two models come from the same lab, and whatever they share in training appears to include this game. OpenAI has not said how GPT-6.1 Sol and GPT-6 Astra are related, so that remains our inference. One overlap points the other way, and there it is only in the names: Opus 5.5 called its level Pinecrest Ridge, Sol called its own Pinecrest Pass, and both open on a section called The Meadows. Sol never opened Opus 5.5’s file either.
Where Sol parts company with Astra, it trades a better picture for better sound. Astra’s buggy is the more detailed drawing, and beside it Sol’s looks plain. Astra’s audio, though, is close to inaudible at a normal volume, an engine held at a gain of 0.026 behind a low-pass filter. Sol’s engine starts at a similar level but rises to 0.047 on the throttle with no filter, and it can be heard. The two also drain their tanks at different rates: 2.05% a second on Astra’s throttle, 0.83% on Sol’s.
Time and size
| Claude Opus 5.5 | GPT-6.1 Sol | |
|---|---|---|
| Lab | Anthropic | OpenAI |
| Harness | Claude Code subagent | Codex CLI |
| Setting | max | max |
| Wall clock | 98 min 6 s | 57 min 38 s |
| Delivered file | 113 KB, 2,072 lines | 58 KB, 194 lines |
| Physics step | 240 Hz, interpolated | 120 Hz |
| Level length | 1,176 m | 900 m |
| Full throttle | Crashes at about 260 m | Rolls at 378 m |
| Fuel | Runs out on a slow driver | Runs out only once the car has stalled |
| Coins | 214, 60 of them in the air | 80, all on the driving line |
GPT-6.1 Sol worked for a little under an hour, 40 minutes less than Opus 5.5, and wrote about half as much. Its 194 lines are long ones: the whole file runs to 58 KB.
Scoring
Opus 5.5 took full marks in graphics, physics and gameplay and was a point short in sound. GPT-6.1 Sol scored 3 for graphics, 4 for physics, 3 for sound and 3 for gameplay.
Graphics and gameplay carry the widest gaps, two points each. Opus 5.5’s world is the fuller one, down to the paint jobs, the rock strata and the driver’s scarf. Sol’s is calm and tidy, with menus that would pass for a magazine layout, around a buggy that looks plain beside Opus 5.5’s and beside Astra’s. In gameplay, Opus 5.5’s level rewards driving well, with coins in the air for the right speed off each lip and a tank that strands a timid driver. Sol’s asks for one thing, keeping the speed down to somewhere between 30 and 50 km/h, and then gives the rest away: every coin comes to anyone who finishes, and the fuel only runs out on a car that has already stopped.
Physics and sound are a point apart. Both cars sit on real springs, shift their weight under throttle and brake, and hold together for a whole run. Opus 5.5’s has more life in it: dirt thrown in proportion to wheelspin, a buggy that coasts when you lift, and jumps long enough to steer in. Sol’s hops last about a second. Opus 5.5 has an engine note that rises with revs and load, coin chimes, landing thuds and a low-fuel alarm, with an on-off switch in the menu and on the pause card. Sol has one engine tone, a chime and a handful of notes, and starts muted.
One brief, two answers
On this brief GPT-6.1 Sol is a long way ahead of the model it replaces. GPT-6 Sol’s 160-metre course could be finished by holding a single key for sixteen seconds. GPT-6.1 Sol’s 900 metres cannot be finished that way, and its buggy rolls if you try. It spent almost twice as long to get there, and what it arrived at was a game its own lab had already made, in layout and in its numbers.
Opus 5.5’s build has now been played beside games from three other models and still stands apart for the same reason. It is the one entry in which the level is part of the challenge: the coins, the fuel and the crashes each ask for a different kind of driving. GPT-6.1 Sol built a game that is pleasant to drive and hard to lose once you have learned to ease off, and 900 metres later it has asked nothing more of you.
How this test was run
One attempt per model, with no follow-up messages and no clean-up: both files are served exactly as they were delivered. GPT-6.1 Sol received the prompt Opus 5.5 was given on 23 September, identical apart from the two lines that name the output folder and file, checked line by line. It ran in OpenAI’s Codex CLI at max, not ultra, which adds automatic task delegation. Opus 5.5 ran as a Claude Code subagent at maximum effort. Both were told to verify their work before reporting done, and both had a shell, Node and a browser to do it with. Wall-clock time depends on the server as well as the model, and the two runs were a week apart.
Scoring was done by BitsMinds across four categories, from a local side-by-side page of the untouched builds with Astra’s Ridgeline open beside them for reference, playing them in an ordinary browser. There was no blind scoring and no automated grading. The measurements quoted above, among them crash distances, coin counts and fuel at fixed speeds, suspension travel and jump lengths, were taken afterwards by driving each delivered build’s own code and reading its own state.
Neither entrant started from a blank page, though nothing on screen shows it. Before writing any code, GPT-6.1 Sol searched its persistent Codex memory for earlier one-shot builds and read what Codex had kept from the Pac-Man round, along with a saved note on single-attempt deliveries. None of it concerns Hill Climb, and Codex entrants did the same in every earlier field on this brief. Claude Code, for its part, handed Opus 5.5 this project’s memory index, a page of one-line notes kept while running the site, which says nothing about this brief, fuel or coins.
Earlier fields on this brief: GPT-6 Astra vs Claude Fable 5.1, Opus 5 vs GPT-5.6 Sol vs Grok 4.7, and Opus 5.5 vs GPT-6 Sol vs Grok 4.7.
More on Claude
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.