Claude Opus 5.5 Review: The New No. 1 Costs Less

Claude Opus 5.5 takes sole first place on the Artificial Analysis Intelligence Index, scoring 58 to the 53 shared by Claude Fable 5.1 and GPT-6 Astra. It lists at $4/$20, a fifth under Opus 5. The better story sits one effort level down: at high it scores 54, above every rival, for $1.82 a task. It also won our three-round build-off, and it was the slowest entrant in every round.

By BitsMindsHands-on reviewPublished Facts checked

The short version

Get Claude if

  • Teams running Claude Opus 5 or Fable 5.1 in production who want the same work for less
  • Agentic coding and long, sprawling jobs — migrations, audits, multi-repo changes
  • Knowledge work where invented figures are unacceptable: reports, financial models, research

Skip it if

  • Latency-sensitive apps that relied on switching thinking off — plan a migration first
  • Workloads where a score in the high 40s is enough: GPT-6 Sol and Muse Spark 1.3 cost far less
  • Browser and business-workflow automation, where GPT-6 Astra is level or ahead

What we tested

  • Three from-scratch build briefs against Fable 5.1 and GPT-6 Astra, one attempt each at maximum effort
  • The Artificial Analysis Intelligence Index v4.3.2 table and model page, read directly on 22 September 2026
  • Anthropic’s launch post, pricing page and API migration notes
Facts checked

Pros

  • First of 212 models on Artificial Analysis Intelligence Index v4.3.2 at 58
  • At high effort scores 54 for $1.82 a task — above every rival at any setting
  • Beats Claude Fable 5.1 by five index points for 22% less per task
  • List price $4/$20 — a fifth under Opus 5 — with cache reads cut 60% to $0.20
  • Terminal-Bench 4.0 at 66.4% against 57.9% for GPT-6 Astra and 55.8% for Fable 5.1
  • Won our three-brief build-off 10–9–8 against Fable 5.1 and GPT-6 Astra
  • Clearer and more direct writing than Opus 5 — the model’s best-known complaint
  • Higher five-hour usage limits on Pro and Max and Team plans at launch

Cons

  • At max effort it emits 260M output tokens on the index — about three times the median
  • Max effort costs $5.98 a task — no cheaper than Opus 5 and nearly twice GPT-6 Astra
  • Slowest entrant in every round of our build-off at 59 to 77 minutes a brief
  • Thinking cannot be switched off — code that disabled it now gets a 400 error
  • Forced tool choice (any or tool) is no longer accepted
  • GPT-6 Astra still wins AutomationBench and Terminal-Bench-Science
  • Safeguards hand some cyber and biology tasks to older Claude models
  • The claimed 40% saving is Anthropic’s own figure and not yet visible on outside bills

Claude Opus 5.5 arrived on 22 September with a claim that usually does not survive independent testing: better than the model it replaces and cheaper to run. Anthropic said it performs at the level of Claude Fable 5.1 on most work. The independent numbers go further. On the Artificial Analysis Intelligence Index it is first of 212 models, with a score of 58, five points clear of Fable 5.1 and GPT-6 Astra.

This is the second time this year a cheaper Claude has taken the top spot. Opus 5 led in July at half of Fable 5’s price. This time the gap is wider: Opus 5.5 lists at $4/$20, and the two models it displaced both list at $10/$50. The per-token price is only part of the verdict, as our Astra review showed, so this review leads with what a finished task costs.

First place, and what it costs

Artificial Analysis runs the same ten evaluations on every model and publishes what one task cost to complete. It measured Opus 5.5 at five effort levels. Four of them are below, and the effort level matters more than anything else about this model:

Model and effortIntelligenceCost per taskList price
Claude Opus 5.5, max58$5.98$4 / $20
Claude Opus 5.5, xhigh56$3.46$4 / $20
Claude Opus 5.5, high54$1.82$4 / $20
Claude Fable 5.1, max53$7.63$10 / $50
GPT-6 Astra, max53$3.26$10 / $50
Claude Opus 5, max51$5.86$5 / $25
Claude Opus 5.5, medium (API default)51$1.34$4 / $20
GPT-6 Sol, max48$1.06
What one benchmark task costs, ordered by scoreUSD per Intelligence Index v4.3.2 task, rounded · score in brackets · lower is better · Artificial AnalysisCost per task (USD)024686.0Opus 5.5 max(58)3.5Opus 5.5 xhigh(56)1.8Opus 5.5 high(54)7.6Fable 5.1 max(53)3.3Astra max (53)5.9Opus 5 max (51)1.3Opus 5.5 medium(51)
Read it left to right. At maximum effort Opus 5.5 is first and costs about what Opus 5 did. One notch down, at high, it still outscores every other model at any setting — for 56% of what GPT-6 Astra spends and under a quarter of Fable 5.1’s bill. Data: Artificial Analysis.

Three readings come out of that table.

At maximum effort it is the best model, and not a cheap one. A score of 58 for $5.98 a task beats Fable 5.1 by five points for 22% less money. It also costs slightly more than Opus 5 did at maximum, and nearly twice what Astra spends. Anthropic’s “40% less to run than Opus 5” is a claim about default settings, and at max it does not hold on this benchmark.

At high effort it is the best value at the top of the market by a distance. A score of 54 is higher than any other model reaches at any setting. It costs $1.82 a task: 56% of Astra’s bill for a point more, and under a quarter of Fable 5.1’s. This is the configuration that justifies first place on value as well as capability.

At medium, the API default, it equals Opus 5 at its best. Both score 51, for $1.34 against $5.86. For a team on Opus 5, that is a 77% cut in cost per task with no drop in score.

One caution about verbosity. At maximum effort Opus 5.5 produced 260 million output tokens across the index, against a median of 88 million, and the full run cost $8,708. That is why max is priced like Opus 5 despite the lower sticker. Opus 5.5 lowers the price per token and then spends more tokens when it is allowed to. The effort setting decides which of those two wins.

Anthropic’s benchmarks, including the ones it loses

The launch table is Anthropic’s own, run in its own harness. It is also unusually candid: it includes GPT-6 Astra, and Astra wins two rows.

Anthropic’s published benchmarksScore, % — higher is better · self-reported by Anthropic; Astra figures as reported by OpenAI · n/r = not reportedClaude Opus 5.5Claude Fable 5.1GPT-6 AstraClaude Opus 502040608066.455.857.952.3Terminal-Bench4.054.450.353.348.0FrontierCodev1.157.851.8n/r46.6CursorBench 4.040.031.441.426.9AutomationBench58.752.664.629.0Terminal-Bench-Science67.765.657.263.6Humanity’s LastExam
Opus 5.5 leads on four of the six rows. GPT-6 Astra takes AutomationBench narrowly and Terminal-Bench-Science clearly. Terminal-Bench is scored at xhigh for Opus 5.5 and high for Astra, each model’s best. Data: Anthropic, OpenAI.

Coding is where the margin is widest. Terminal-Bench 4.0 is the benchmark closest to an agent working in a real shell, and Opus 5.5 scores 66.4% there against Astra’s 57.9% and Fable 5.1’s 55.8%. Anthropic adds a cost claim for FrontierCode: at default effort Opus 5.5 scores 54.6%, above Astra’s best, for about a fifth of the cost per task. That is Anthropic’s figure, but it points the same way as the independent cost table above.

Astra keeps two categories. On AutomationBench, Zapier’s test of business workflows across connected apps, it scores 41.4% to Opus 5.5’s 40.0%. On Terminal-Bench-Science it leads 64.6% to 58.7%. Anthropic also notes that its safeguards intervened on some cyber and biology tasks, and that another Claude model completed those tasks. That probably lowered Opus 5.5’s scores, and it is also how the model will behave in production.

GDPval-AA v2.1, Elo points against Claude Opus 5Real work across 44 occupations · Opus 5 = 1708 · figures published by Anthropicahead of Opus 5behind Opus 5-2000+200Claude Opus 5.5+138Claude Fable 5.1+27GPT-5.6 Sol−120GPT-6 Astra−166
The widest gap in Anthropic’s table is on professional work, not code: 111 Elo over Fable 5.1 and roughly 300 over GPT-6 Astra. Artificial Analysis runs GDPval-AA itself, but these are the figures Anthropic published.

The largest gap is on professional work rather than code. On GDPval-AA v2.1, real tasks across 44 occupations, Opus 5.5 is 138 Elo above Opus 5 and 111 above Fable 5.1. Anthropic reports a related internal test in which Opus 5.5 wrote company reports from a web copy where the earnings release was hard to find, and 16 of its 18 reports cleared a quality bar that any invented figure or quote would fail. Fable 5.1 and Opus 5 did not clear it once. That test is Anthropic’s own. We would still take it over any leaderboard when deciding what to trust with a financial model.

What we found running it ourselves

We gave Opus 5.5 the same three build briefs that Fable 5.1 and GPT-6 Astra had already attempted: a motorway interchange and a seaside fairground as animated SVGs, and a cinematic rocket launch as one web page. Each model had one attempt at maximum effort and could not look at what it had built. Opus 5.5 won 10–9–8.

It is the best planner we have tested. Its carousel is the first in the series to behave like a solid object in perspective: the horses’ orbit matches the roof’s proportions, and the horses pass behind the centre pole and in front of it at the right moments. Its motorway merges were timed so that no car taking a ramp came within about 145 pixels of other traffic. Its rocket flies a complete mission to a 203 by 207 km orbit, and it stops accelerating at the speed that orbit actually requires. Neither rival’s rocket ever shut its engines down.

It is also slow. It took 59, 63 and 77 minutes on the three briefs, more than twice Astra’s time on each, and nearly all of that was spent thinking. On the interchange it reasoned for 57 minutes and then wrote the file in a minute and a half. Its one real fault was something thinking cannot catch: painted edge lines that run straight across every ramp junction. The fault is only visible in the picture, and the test did not let any model look at its output.

That matches the independent profile: very high capability, very high token use at maximum effort, and results that justify the wait when the work needs planning. It is a model to hand a long task and leave alone. It is not a good choice for a quick question at max effort.

What changes for developers

Opus 5.5 is not a drop-in replacement. Anthropic’s migration notes list four breaking changes from Opus 5:

  • Thinking is always on. A request that disables thinking, or sets a manual token budget, returns a 400 error. Effort (low, medium, high, xhigh, max) is now the only control, and the API defaults to medium.
  • Forced tool use is gone. tool_choice set to any or a named tool now returns an error. The replacement is auto with strict tool use.
  • Thinking blocks are tied to the model and the conversation, so they cannot be carried across to another model mid-thread.
  • The older computer_20251124 computer-use tool is not accepted on the Claude API or Google Cloud.

One further change breaks no request but can make an app go quiet: text the model writes between tool calls now arrives inside thinking blocks, empty at the default display setting. An app that streams that text as progress updates needs to change the display setting.

The price sheet itself is simple. It is $4 in and $20 out per million tokens. Cache reads cost $0.20, 5% of the input price rather than the usual 10%, and five-minute cache writes $5. Batch halves that to $2/$10. Fast mode, up to 2.5 times the speed, costs $8/$40. The pricing calculator now carries Opus 5.5 alongside the rest of the field. On subscriptions, Anthropic raised five-hour usage limits on Pro, Max, Team and seat-based Enterprise plans at launch. Our Claude plans review covers what each tier includes.

Safety is part of the product this time. Anthropic rates Opus 5.5 comparable to Claude Mythos 5.1 in biology and cybersecurity and ships it with Fable 5.1’s safeguards, which is why Artificial Analysis lists the configuration “with fallback”. Vetted organisations can apply for fuller biology access now, and a cyber programme is expanding in the coming weeks. For most work none of this is visible. For security research it matters.

Who should be running it

If you are…Do this
On Claude Opus 5Move. It is better on every row Anthropic published, cheaper per token, and at medium it matches Opus 5’s max score for under a quarter of the cost. Budget a day for the breaking API changes
On Claude Fable 5.1Re-run your evals. Opus 5.5 scores higher for less than half the per-token price; Fable 5.1 is still the model to keep if your own results favour it
On GPT-6 AstraTest Opus 5.5 at high on coding and knowledge work, where it wins on both score and cost. Keep Astra for computer use and business-workflow automation
Choosing an effort levelStart at high, not max. Max buys four index points for more than three times the cost per task
Serving interactive usersStay at medium or high. Artificial Analysis measured 165 seconds to the first output at xhigh
Happy with a score in the high 40sGPT-6 Sol and Muse Spark 1.3 finish a task for $1.06 and $1.60. Opus 5.5 is more model than you need

Verdict

Opus 5.5 earns the top score on this site. That is a 4.9 out of 5, and no review here has scored higher. It is the most capable model on the independent index. Its high setting outscores every rival for less than any of them charge at the top. It won our hands-on build-off, and it did that by thinking problems through rather than by decoration.

It misses the last tenth for reasons that are specific and manageable. At maximum effort it is verbose and slow, so the cheaper sticker does not produce a cheaper task. Removing the option to disable thinking is a real migration cost for anyone already in production. Astra still leads on computer use and business workflows. And the 40% saving Anthropic promises will not be confirmed until customers see it on their own bills.

Our advice is to run it at high, keep max for work that justifies a long wait, and if you have not re-run your evaluations since Opus 5.5 shipped, assume you are overpaying. Sonnet 5.5 and Haiku 5.5 are due in the coming weeks, and the leaderboard will be updated when they are measured.

About the score. 4.9/ 5 is BitsMinds' editorial verdict from our own testing and research — not an average of user ratings, which we do not collect. Prices and plan tiers are as published by the vendor on the fact-check date shown above. How we rate →

Related Reviews

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.