Claude Opus 5.5 Is Here: Cheaper Than Opus 5, and Better
Anthropic shipped Claude Opus 5.5 on 22 September at $4 and $20 per million tokens, a fifth under Opus 5 and ahead of it on every benchmark it published. The sticker fell 20% but Anthropic claims a task costs 40% less, because the model also spends fewer tokens getting there. Every number here is Anthropic's own.
Claude Opus 5.5 is Anthropic's new flagship Opus model, announced on 22 September 2026 and available immediately. It answers to the model ID claude-opus-5-5, carries a 1M-token context window with up to 128K tokens of output, and costs $4 per million input tokens and $20 per million output — a fifth less than Claude Opus 5 on both headline rates. It is live on the Claude API, in Claude Code, and on AWS, Google Cloud and Microsoft Azure.
Anthropic's own summary is that it "performs at the level of Claude Fable 5.1 on most work" while costing substantially less to run than Opus 5. The benchmark tables it published go further than that in coding, where Opus 5.5 is ahead of Fable 5.1 by a comfortable margin rather than level with it.
The price cut is bigger than it looks
The sticker is the smaller half of the story. On the pricing page Opus 5.5 lists at $4 in and $20 out against Opus 5's $5 and $25 — a flat 20% cut. Cache reads fall further, from $0.50 to $0.20, and five-minute cache writes from $6.25 to $5.
Anthropic's claim, though, is that a typical job costs 40% less, not 20%. The difference is token count: the company says Opus 5.5 uses about 20% fewer input and output tokens to reach the same place. A fifth off the rate and a fifth fewer tokens compound to roughly a third to two fifths off the invoice, which is why the per-token number and the per-task number disagree.
That distinction matters more than it sounds. Two models with identical sticker prices routinely differ by a wide margin on the bill, because the one that reasons in fewer tokens simply buys less of them. Anthropic also says output generates more than 30% faster than Opus 5, which lands on wall-clock time rather than cost.
Where the benchmarks actually separate
Every figure that follows is Anthropic's own, published alongside the model and run with adaptive thinking at maximum effort. None of it has been independently reproduced yet, here or anywhere else.
Coding is where the model makes its case. Terminal-Bench 4.0 is the standout at 66.4% against Fable 5.1's 55.8% and Opus 5's 52.3% — a 14-point jump over the model it replaces, on the benchmark that most resembles an agent working in a real shell.
Away from code the picture is flatter. On Humanity's Last Exam, OSWorld 2.0 and Chartography, Opus 5.5 clears Fable 5.1 by between 0.6 and 2.1 points — close enough that "performs at the level of Fable 5.1" is the honest description. The gap that stays wide in every one of them is against Opus 5.
AutomationBench is the outlier. Anthropic puts Opus 5.5 at 40.0% against 28.8% for GPT-5.6 Sol and 26.9% for Opus 5, and published no Fable 5.1 figure for it at all. A missing number in a vendor's own table is worth noticing, though there are innocent reasons for one.
Thinking is no longer optional
One behavioural change lands on anyone with code in production: switching thinking off is no longer available. Opus 5.5 thinks adaptively, always, deciding for itself how much reasoning a request warrants. Applications that explicitly disabled thinking for latency or cost will need to be revisited.
Anthropic has also applied an anti-distillation measure it calls preserved thinking to API accounts created after 31 August 2026. The practical effect is on what those accounts can read back out of the reasoning trace.
The leak got this one right
The release arrived exactly as it had been described. We covered the rumour on 21 September and were sceptical of it, because the same chain had already called the model Fable 5.2 and then Opus 5.2, and because nothing in it was checkable: no model ID, no documentation page, no pricing row. The scepticism was about the evidence, and the evidence was genuinely thin. The forecast was still wrong. The leaked sheet named $4 and $20, cache reads at $0.20, cache writes at $5 and a Tuesday launch, and every one of those held.
What it could not produce in advance was the only thing that ever settles it, and that is the part worth keeping. A model ID, a docs entry and a pricing line appeared on 22 September, and until they did, a price sheet circulating on one account was still just a price sheet.
The number to watch now is the one Anthropic put at the centre of its own pitch. A 20% price cut is easy to verify and already confirmed on the pricing page. A 40% drop in what a real task costs depends on a token-efficiency claim that only shows up on an actual bill, over actual work, and nobody outside Anthropic has one of those yet.
Coming next: Opus 5.5 goes into the Lab against Claude Fable 5.1 and GPT-6 Astra — the same briefs every frontier model in this series has been handed, one attempt each, at maximum reasoning effort, with every build playable in the article. The September round put those two against Grok 4.7; this one adds the model above. It will land in the Lab.
More on Claude
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.