BitsMinds Lab

Our own tests, not a summary of someone else's benchmark. We hand frontier models the same build brief, take one attempt each, and publish every output untouched so you can judge it yourself.

How the tests work

One brief, one attempt
Every model gets the identical brief. One run each — no retries, no follow-up prompts, no clean-up of the output.
Published untouched
The file each model wrote is served as-is and runs live inside the article, so the evidence is the artefact, not our description of it.
Scored in the open
Each round awards points for finishing position (3-2-1 in season two, 4-3-2-1 from season three); ties are shared. The scoring is stated on every article.
One judge
A single BitsMinds editor scores the outputs from a local side-by-side page. It is not blind and not a panel — read it as one informed opinion with the evidence attached.
No browser for the model
Models cannot open what they built to check it. That constraint is deliberate: it is what separates models that re-read their own work from those that do not.
What was not recorded
Reasoning-effort settings were not logged for the first two seasons; from season three every model ran at maximum effort. Each article's conditions panel says exactly what is known for that run.

These are single runs. A different attempt could land differently; treat the pattern across rounds, not any one score, as the finding.

All tests, newest first

Astra 6 Fable 5.1 VS BITSMINDS.COM
Season 3

GPT-6 Astra vs Claude Fable 5.1: There Can Be Only One

OpenAI’s GPT-6 Astra and Anthropic’s Claude Fable 5.1 got the same three build briefs — a motorway interchange, a seaside fairground and a cinematic rocket launch — one attempt each, at maximum reasoning effort, and neither could open a browser to check its own work. All six builds run live inside the article, faults and all: a bridge that hides a truck mid-crossing, a carousel that is not doing what it looks like it is doing, and two rockets with very different ideas of what a launch is. Score it yourself.

Models
GPT-6 Astra, Claude Fable 5.1
Briefs
3 · one attempt per model per brief
Effort
Maximum reasoning effort for both models
Result
GPT-6 Astra 9 pts · Fable 5.1 10 pts
Read the test →
Fable 5.1 17 Opus 5 16 Fable 5 13 Sonnet 5 6 BITSMINDS.COM
Season 3

Claude Fable 5.1 vs Opus 5 vs Sonnet 5: One Point Apart

We gave four Claude models — Fable 5.1, Opus 5, Fable 5 and Sonnet 5 — the same five build briefs at maximum reasoning effort, one attempt each, with no browser to check their work. All twenty builds run live inside the article. Fable 5.1 won four of the five rounds and took the season by a single point, because the round it lost was the one where its code never ran at all.

Models
Claude Fable 5.1, Claude Opus 5, Claude Fable 5, Claude Sonnet 5
Briefs
5 · one attempt per model per brief
Effort
Maximum reasoning effort for every model
Result
Fable 5.1 17 pts · Opus 5 16 pts · Fable 5 13 pts · Sonnet 5 6 pts
Read the test →
8 Opus 5 6 Fable 5 3 Sonnet 5 BITSMINDS.COM
Season 2

Claude Opus 5 vs Fable 5 vs Sonnet 5: One Prompt Each

We gave Claude Opus 5, Fable 5 and Sonnet 5 the same three build briefs — an animated SVG fairground, a falling-sand physics sandbox and a self-solving Rubik’s cube — one run each, no retries, no browser. All nine builds run live inside the article. Opus 5 takes it 8–6–3, and the biggest surprise is the clock: the largest model was the fastest on every round, by a factor of nearly three.

Models
Claude Opus 5, Claude Fable 5, Claude Sonnet 5
Briefs
3 · one run per model per brief
Effort
Not recorded for this run
Result
Opus 5 8 pts · Fable 5 6 pts · Sonnet 5 3 pts
Read the test →
BITSMINDS.COM
Season 1

Fable 5 vs Opus 4.8, One Prompt Each: Four Builds, Dead Level at 2–2

We pasted identical prompts into Claude Opus 4.8 and Claude Fable 5 — an animated highway interchange, a one-file Space Invaders, a pure-SVG aquarium, and a 3D rocket launch — one shot each, no retries. Everything runs live inside the article. Four builds in, it is dead level at 2–2 — and the launch round flips the very pattern the first three revealed.

Models
Claude Opus 4.8, Claude Fable 5
Briefs
3 · one attempt per model per brief
Effort
Not recorded for this run
Result
Opus 4.8 2 rounds · Fable 5 2 rounds
Read the test →

Want the numbers rather than the pictures? The model leaderboard carries the independent Artificial Analysis scores, and our reviews cover the same models as products.