Claude Sonnet 5.5 Review: Great Value Until You Max It

Claude Sonnet 5.5 scores 56 on the Artificial Analysis Intelligence Index, second only to Claude Opus 5.5, at Sonnet 5's unchanged $2/$10. Up to high effort it is cheap and fast, answering in seconds where Opus 5.5 takes half a minute. At maximum effort it costs more per task than Opus 5.5 for a lower score.

By BitsMindsHands-on reviewPublished Facts checked

The short version

Get Claude if

  • Teams on Claude Sonnet 5, for whom it is a same-price upgrade
  • Chat and interactive products, where it answers in seconds at medium effort
  • Everyday coding, bug fixing and documents where high effort is enough

Skip it if

  • Anyone planning to run it at max effort: Opus 5.5 reaches the same score for less than half the cost
  • Work that needs the top score: Opus 5.5 is ahead at every setting
  • Very cheap bulk jobs where a score in the high 30s is enough: GPT-6 Luna costs about a sixth as much

What we tested

  • The Artificial Analysis Intelligence Index v4.3.2 table and all five Sonnet 5.5 effort settings, read directly on 29 September 2026
  • Anthropic’s launch post, model page, pricing and migration guide
Facts checked

Pros

  • Second on the Artificial Analysis Intelligence Index at 56 — behind only Claude Opus 5.5
  • Same $2/$10 list price as Sonnet 5 with cache reads at $0.20
  • At medium effort beats Sonnet 5’s best score for about a ninth of the cost per task
  • At high effort scores 47 for $1.08 a task — a point behind GPT-6 Sol’s best at almost the same cost
  • Answers in about 6 seconds end to end at medium where Opus 5.5 needs 18 to 29
  • Terminal-Bench 4.0 at 70.6% — ahead of Opus 5.5’s best of 66.4%
  • Computer use up from 57.0% to 80.1% on OSWorld 2.1 in Anthropic’s tests
  • On every major cloud from launch day with zero data retention available

Cons

  • Max effort costs $7.60 a task — more than Opus 5.5 at max for a lower score
  • Opus 5.5 at xhigh reaches the same score of 56 for $3.46
  • Generated 410M output tokens on the index at max — nearly five times the median
  • Opus 5.5 at low scores a point higher than Sonnet 5.5 at medium for less money
  • Thinking cannot be disabled — between_tools is the floor and only up to high effort
  • Forced tool choice and the old computer-use tool now return errors
  • Higher-risk security tasks fall back to Sonnet 5
  • Edited conversation histories can invalidate its thinking blocks on newer accounts

Claude Sonnet 5.5 arrived on 28 September at exactly Sonnet 5’s price, $2 in and $10 out per million tokens, with benchmark scores Anthropic put close to Opus 5.5. The independent numbers confirm most of that. On the Artificial Analysis Intelligence Index it scores 56 at maximum effort, second only to Opus 5.5’s 58, and ahead of Claude Fable 5.1 and GPT-6 Astra at 53. Sonnet 5 scored 38 on the same index, so the gain over the model it replaces is 18 points.

Our Opus 5.5 review found that the effort setting decides the verdict more than the model does. That is even more true of Sonnet 5.5, because here the answer changes sides: at some settings it is the best value Anthropic sells, and at others its own sibling beats it on both score and price.

Second on the index, and what each score costs

Artificial Analysis ran Sonnet 5.5 at all five effort settings and publishes what one task cost to complete. It also times the complete response, which it defines as input processing, reasoning and answer generation together. Here is every Sonnet 5.5 and Opus 5.5 setting, ordered by score, with the models they are usually compared against:

Model and effortIntelligenceCost per taskFull responseList price
Claude Opus 5.5, max58$5.98685 s$4 / $20
Claude Sonnet 5.5, max56$7.60374 s$2 / $10
Claude Opus 5.5, xhigh56$3.46150 s$4 / $20
Claude Opus 5.5, high54$1.8262 s$4 / $20
GPT-6 Astra, max53$3.26314 s$10 / $50
Claude Sonnet 5.5, xhigh52$2.7458 s$2 / $10
Claude Opus 5.5, medium51$1.3428.5 s$4 / $20
GPT-6 Sol, max48$1.05186 s$2 / $10
Claude Sonnet 5.5, high (API default)47$1.0822.6 s$2 / $10
Claude Opus 5.5, low42$0.5518.5 s$4 / $20
Claude Sonnet 5.5, medium (apps default)41$0.596.0 s$2 / $10
Claude Sonnet 5, max38$5.09190 s$2 / $10
Claude Sonnet 5.5, low36$0.416.8 s$2 / $10
Intelligence Index at each effort settingArtificial Analysis Intelligence Index v4.3.2, 0–100, axis from zero · read 29 September 2026Claude Sonnet 5.5Claude Opus 5.50102030405060703642Low4151Medium4754High5256Xhigh5658Max
Opus 5.5 is ahead at every setting, by ten points at Medium and two at Max. Sonnet 5 scored 38 at its maximum on the same index. Data: Artificial Analysis.
What one index task costs at each effort settingUS cents — lower is better · Artificial Analysis, index v4.3.2Claude Sonnet 5.5Claude Opus 5.502004006008004155Low59134Medium108182High274346Xhigh760598Max
Sonnet 5.5 is the cheaper model at every setting except Max, where it spends 410 million output tokens on the index to Opus 5.5’s 260 million.

Three readings come out of those numbers.

Up to High, it is excellent value. Medium, the default in Claude’s apps and Claude Code, scores 41 for $0.59 a task. That already beats Sonnet 5’s best score of 38, which cost $5.09, for about a ninth of the price. High, the default on the API, scores 47 for $1.08. That is a point behind GPT-6 Sol at its maximum at almost exactly the same cost, and Sonnet 5.5 has far more headroom above it.

Above High, the value belongs to Opus 5.5. Sonnet 5.5 at Xhigh scores 52 for $2.74, while Opus 5.5 at High scores 54 for $1.82. At Max it reaches 56 for $7.60, and Opus 5.5 reaches the same 56 at Xhigh for $3.46, less than half as much. The reason is token use. At Max, Sonnet 5.5 generated 410 million output tokens across the index, against a median of 88 million and 260 million for Opus 5.5 at its own maximum. A cheaper price per token buys nothing when the model spends more tokens on the same task.

At the bottom of the range, the two overlap. Opus 5.5 at Low scores 42 for $0.55, a point above Sonnet 5.5 at Medium for four cents less. On score and cost alone, Sonnet 5.5 is the cheapest route to a given score only at its Low and High settings. What settles the overlap is the column the index does not score.

Speed is the reason to choose it

Sonnet 5.5 is much faster. At Medium it returns a complete answer in about 6 seconds, with the first output after 1.2 seconds. Opus 5.5 takes 18.5 seconds at Low and 28.5 at Medium. At High the gap is 22.6 seconds against 62.2, and at Max it is about 6 minutes against 11. Output also streams faster at every setting: 85 to 139 tokens a second across Sonnet 5.5’s range, against 74 to 93 for Opus 5.5.

For anything a person waits on, such as a chat reply, a code suggestion or a support answer, that difference outweighs a few cents per task. It is the case for Sonnet 5.5 even at the points where Opus 5.5 is cheaper. Anthropic’s partners report the same pattern in their own tests, though Anthropic chose which to publish. Balyasny Asset Management measured about 121,000 tokens per answer against Sonnet 5’s 497,000 on 2,441 finance tasks, Zendesk says tickets were processed 20% faster, and Box reports 2.4 times the speed with 12% fewer tokens.

Anthropic’s benchmarks

The launch table is Anthropic’s own, run in its own harness. It puts Sonnet 5.5 a few points under Opus 5.5 almost everywhere, with one exception.

Anthropic’s published benchmarksScore, % — higher is better · self-reported by Anthropic, 28 September 2026Claude Sonnet 5.5Claude Opus 5.5Claude Sonnet 502040608010070.666.410.3Terminal-Bench4.052.154.442.4FrontierCodev1.155.557.834.1CursorBench 4.064.567.754.9Humanity’s LastExam80.181.857.0OSWorld 2.161.664.415.6Chartography
Terminal-Bench is the only row Sonnet 5.5 wins, and its FrontierCode bar is its best score, at Xhigh; at Max it drops to 46.2%. HLE is scored with tools, OSWorld with partial credit, Chartography without tools. Data: Anthropic.

That exception is Terminal-Bench 4.0, where Sonnet 5.5 scores 70.6% against Opus 5.5’s best of 66.4%. Anthropic lists Sonnet 5 at 10.3% on the same test and does not explain that figure. On FrontierCode, Anthropic flags a fault of its own: at Max, Sonnet 5.5 more often ran Claude Code’s multi-agent review skill and changed things outside the task, and it scored 46.2% there against 52.1% at Xhigh. Leaving Terminal-Bench aside, the largest gains over Sonnet 5 are in computer use and chart reading, which rise by 23 and 46 points.

Professional work, Elo points above Sonnet 5Sonnet 5 = 0 (1449 on GDPval-AA, 1359 on AA-Briefcase) · run by Artificial Analysis, published by AnthropicClaude Sonnet 5.5Claude Opus 5.5GPT-6 Sol010020030040050039539738GDPval-AA v2.1452463124AA-Briefcasev1.1
On both boards Sonnet 5.5 finishes within 11 Elo points of Opus 5.5. OpenAI has since fixed an image-understanding bug in GPT-6 Sol, so its two scores may not reflect the current model.

Professional work is where Sonnet 5.5 comes closest to Opus 5.5. On GDPval-AA, which covers real tasks from 44 occupations, it finishes two Elo points behind, and on AA-Briefcase eleven. Artificial Analysis ran both boards on a pre-release deployment with a structured-output bug that has since been fixed. Anthropic expects any effect to be small and to understate Sonnet 5.5.

What changes for developers

The price list is identical to Sonnet 5’s: $2.50 per million tokens for five-minute cache writes, $4 for one-hour writes, $0.20 for cache reads, and half price on the Batch API. The tokenizer is also unchanged. The API behaviour is not. Anthropic lists five breaking changes:

  • Thinking can be reduced but not disabled. thinking: {"type": "disabled"} returns a 400 error. The floor is a new between_tools mode, which skips up-front thinking and is accepted only up to High effort.
  • Forced tool use is gone, as it is on Opus 5.5. Use auto with strict tool use.
  • Thinking blocks are tied to the model and the conversation. On accounts created on or after 31 August 2026, replaying one after editing an earlier turn is an error, so keep histories append-only.
  • Computer use needs the new toolset on the Claude API and Google Cloud.
  • The advisor tool no longer pairs a Sonnet 5.5 executor with Opus 4.8, Opus 4.7 or Sonnet 5.

One further change can make an app go quiet: text written between tool calls now arrives inside thinking blocks. Sonnet 5.5 is also the first Sonnet with the cyber safeguards Anthropic uses on its top models, so higher-risk security tasks visibly fall back to Sonnet 5. For ordinary development none of this shows. For security research it does.

Who should be running it

If you are…Do this
On Claude Sonnet 5Move. Same list price, 18 more index points at max, and at medium it beats Sonnet 5’s best for about a ninth of the cost per task. Budget time for the breaking API changes
Building something people wait onUse it at medium. It returns a complete answer in about 6 seconds, three to five times faster than Opus 5.5 at low or medium
Running long agentic jobsUse high, the API default. If you need more, switch to Opus 5.5 rather than raising the effort: Opus 5.5 at high scores 54 for $1.82, and at xhigh matches Sonnet 5.5’s max score for less than half the cost
On Claude Opus 5.5Keep it for anything above high effort. Try Sonnet 5.5 where response time matters more than the last few index points
On GPT-6 SolAt high, Sonnet 5.5 is a point behind Sol’s best score at the same cost per task, and answers in 22.6 seconds against 186. Above that, it has headroom Sol does not
Running cheap bulk jobsIf a score in the high 30s is enough, GPT-6 Luna at max costs $0.07 a task against $0.41 for Sonnet 5.5 at low

Verdict

Sonnet 5.5 earns 4.5 out of 5. It is the second-best model on the independent index at the same list price as the model it replaces. It is the fastest way to a strong answer from Claude, and from Low to High effort it is a genuine bargain.

It stops short of Opus 5.5’s 4.9 for one reason, and that reason is Anthropic’s own flagship. Above High, Opus 5.5 reaches every score Sonnet 5.5 can for less money, and at Max the gap is more than double. The breaking API changes are a smaller cost, the same kind Opus 5.5 brought, and worth a day of migration work before switching.

The practical rule fits in one line: Medium for anything a person is waiting on, High for agentic work, and when that is not enough, Opus 5.5 rather than Max. The leaderboard now carries it.

About the score. 4.5/ 5 is BitsMinds' editorial verdict from our own testing and research — not an average of user ratings, which we do not collect. Prices and plan tiers are as published by the vendor on the fact-check date shown above. How we rate →

Related Reviews

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.