Models·6 min read
By BitsMindsSource: Anthropic

Claude Opus 5.5 Is Here: Cheaper Than Opus 5, and Better

Anthropic shipped Claude Opus 5.5 on 22 September at $4 and $20 per million tokens, a fifth under Opus 5 and ahead of it on every benchmark it published. The sticker fell 20% but Anthropic claims a task costs 40% less, because the model also spends fewer tokens getting there. Every number here is Anthropic's own.

Claude Opus 5.5 — a new chapter A monumental copper Claude starburst is suspended inside a sculpted ivory paper ring. Dozens of finely separated pages rise from a stepped circular plinth, then converge into one smooth copper-edged ribbon. Warm light catches the paper layers and brushed copper, casting a long architectural shadow. Large dark typography reads Claude Opus 5.5, with the confirmed input and output list prices and one-million-token context below. The sculpture is an editorial metaphor for reasoning and efficiency, not a diagram or benchmark claim. Original self-contained vector artwork for https://www.bitsminds.com/news/claude-opus-5-5-launch-price-benchmarks-2026. Release: 22 September 2026. Price text reflects the article's standard API list prices, USD per million tokens. Official Claude outline preserved verbatim from the project logo. No raster assets, animation, filters or external dependencies. ANTHROPIC Claude Opus 5.5 $4 in · $20 out per million tokens 1M context BITSMINDS.COM
Share:

Claude Opus 5.5 is Anthropic's new flagship Opus model, announced on 22 September 2026 and available immediately. It answers to the model ID claude-opus-5-5, carries a 1M-token context window with up to 128K tokens of output, and costs $4 per million input tokens and $20 per million output — a fifth less than Claude Opus 5 on both headline rates. It is live on the Claude API, in Claude Code, and on AWS, Google Cloud and Microsoft Azure.

Anthropic's own summary is that it "performs at the level of Claude Fable 5.1 on most work" while costing substantially less to run than Opus 5. The benchmark tables it published go further than that in coding, where Opus 5.5 is ahead of Fable 5.1 by a comfortable margin rather than level with it.

The price cut is bigger than it looks

The sticker is the smaller half of the story. On the pricing page Opus 5.5 lists at $4 in and $20 out against Opus 5's $5 and $25 — a flat 20% cut. Cache reads fall further, from $0.50 to $0.20, and five-minute cache writes from $6.25 to $5.

Anthropic's claim, though, is that a typical job costs 40% less, not 20%. The difference is token count: the company says Opus 5.5 uses about 20% fewer input and output tokens to reach the same place. A fifth off the rate and a fifth fewer tokens compound to roughly a third to two fifths off the invoice, which is why the per-token number and the per-task number disagree.

That distinction matters more than it sounds. Two models with identical sticker prices routinely differ by a wide margin on the bill, because the one that reasons in fewer tokens simply buys less of them. Anthropic also says output generates more than 30% faster than Opus 5, which lands on wall-clock time rather than cost.

Where the benchmarks actually separate

Every figure that follows is Anthropic's own, published alongside the model and run with adaptive thinking at maximum effort. None of it has been independently reproduced yet, here or anywhere else.

Coding and terminal workAccuracy, % — higher is better · Anthropic's own figures, 22 September 2026Claude Opus 5.5Claude Fable 5.1Claude Opus 502040608010066.455.852.3Terminal-Bench4.054.450.348.0FrontierCodev1.157.851.846.6CursorBench 4.0
The widest margin is Terminal-Bench, where Opus 5.5 lands 10.6 points above Fable 5.1 and 14.1 above Opus 5. Source: Anthropic.

Coding is where the model makes its case. Terminal-Bench 4.0 is the standout at 66.4% against Fable 5.1's 55.8% and Opus 5's 52.3% — a 14-point jump over the model it replaces, on the benchmark that most resembles an agent working in a real shell.

Knowledge and computer useAccuracy, % — HLE and Chartography with tools, OSWorld 2.0 partial · Anthropic's own figuresClaude Opus 5.5Claude Fable 5.1Claude Opus 502040608010067.765.663.6Humanity's LastExam81.880.774.0OSWorld 2.089.088.483.4Chartography
Outside coding the gap narrows to roughly a point over Fable 5.1 — these are the benchmarks where Anthropic's “level of Fable 5.1” claim is doing the work.

Away from code the picture is flatter. On Humanity's Last Exam, OSWorld 2.0 and Chartography, Opus 5.5 clears Fable 5.1 by between 0.6 and 2.1 points — close enough that "performs at the level of Fable 5.1" is the honest description. The gap that stays wide in every one of them is against Opus 5.

AutomationBenchAccuracy, % — higher is better · Anthropic did not publish a Fable 5.1 figureClaude Opus 5.5GPT-5.6 SolClaude Opus 50102030405040.028.826.9AutomationBench
The largest relative gap Anthropic published: roughly half again the score of either comparison model, on the benchmark closest to long-running autonomous work.

AutomationBench is the outlier. Anthropic puts Opus 5.5 at 40.0% against 28.8% for GPT-5.6 Sol and 26.9% for Opus 5, and published no Fable 5.1 figure for it at all. A missing number in a vendor's own table is worth noticing, though there are innocent reasons for one.

GDPval-AA v2.1, Elo points against Opus 5Opus 5 = 1708 Elo, set as the zero line · an Elo scale cannot be drawn from zeroahead of Opus 5behind Opus 5-1600+160Claude Opus 5.5+138Claude Fable 5.1+27
On Anthropic's professional-task Elo, Opus 5.5 is 138 points above Opus 5 and 111 above Fable 5.1 — the one board where it does not merely match Fable.

Thinking is no longer optional

One behavioural change lands on anyone with code in production: switching thinking off is no longer available. Opus 5.5 thinks adaptively, always, deciding for itself how much reasoning a request warrants. Applications that explicitly disabled thinking for latency or cost will need to be revisited.

Anthropic has also applied an anti-distillation measure it calls preserved thinking to API accounts created after 31 August 2026. The practical effect is on what those accounts can read back out of the reasoning trace.

The leak got this one right

The release arrived exactly as it had been described. We covered the rumour on 21 September and were sceptical of it, because the same chain had already called the model Fable 5.2 and then Opus 5.2, and because nothing in it was checkable: no model ID, no documentation page, no pricing row. The scepticism was about the evidence, and the evidence was genuinely thin. The forecast was still wrong. The leaked sheet named $4 and $20, cache reads at $0.20, cache writes at $5 and a Tuesday launch, and every one of those held.

What it could not produce in advance was the only thing that ever settles it, and that is the part worth keeping. A model ID, a docs entry and a pricing line appeared on 22 September, and until they did, a price sheet circulating on one account was still just a price sheet.

The number to watch now is the one Anthropic put at the centre of its own pitch. A 20% price cut is easy to verify and already confirmed on the pricing page. A 40% drop in what a real task costs depends on a token-efficiency claim that only shows up on an actual bill, over actual work, and nobody outside Anthropic has one of those yet.

Coming next: Opus 5.5 goes into the Lab against Claude Fable 5.1 and GPT-6 Astra — the same briefs every frontier model in this series has been handed, one attempt each, at maximum reasoning effort, with every build playable in the article. The September round put those two against Grok 4.7; this one adds the model above. It will land in the Lab.

More on Claude

Evergreen coverage we keep current — start here.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Hill Climb: three labs, three little worlds Three miniature cross-sections of game terrain sit side by side on an ivory field. An orange open-top jeep with a helmeted driver climbs a sunset ridge for Claude Opus 5. An orange rover with its small cab below the chassis crosses a moonlit pine forest for GPT-5.6 Sol. A teal number-47 buggy drives through a canyon for Grok 4.7. Coins, star badges, suspension springs and exposed layers of rock make each scene feel like a small working model. The official model-provider marks and model names appear above. Original editorial illustration inspired by the article's screenshots, without scores or simulated rankings. Original SVG illustration for https://www.bitsminds.com/news/claude-opus-5-vs-gpt-5-6-sol-vs-grok-4-7-hill-climb-2026. Three game builds are the subject, not benchmark results. All artwork is vector; brand paths retained from project assets. 47 Opus 5 GPT-5.6 Sol Grok 4.7 BITSMINDS.COM HILL CLIMB
Models

Opus 5 vs GPT-5.6 Sol vs Grok 4.7: Can They Build a Game?

Grok 4.7, Fable 5.1 and Astra 6 — can it compete? An editorial model-making bench holds a miniature motorway interchange, a copper Ferris wheel and a rocket travelling outside a camera frame. A thick clay-orange cable ends in two separated connectors in the foreground, representing the gap between working mechanisms and visible results in Grok's builds. The words Can it compete? sit to the left. The official Grok, Claude and OpenAI marks label the three entrants along the bottom. The miniatures are original illustrations, not screenshots or quantitative comparisons. Original vector editorial illustration for BitsMinds Lab. Article: https://www.bitsminds.com/news/grok-4-7-vs-fable-5-1-vs-astra-6-build-off-2026. Based on the supplied article and published comparison. No benchmark scores are encoded in the composition. CAN IT COMPETE? Grok 4.7 Fable 5.1 Astra 6 VS VS BITSMINDS.COM
Models

Grok 4.7 vs Fable 5.1 vs Astra 6: Can It Compete?

SPACEXAI Grok 4.7 $2 IN · $6 OUT PER MILLION TOKENS BITSMINDS.COM
Models

Grok 4.7 Chases Fable 5.1 at a Fraction of the Price