Claude Sonnet 5.5 Review: Great Value Until You Max It
Claude Sonnet 5.5 scores 56 on the Artificial Analysis Intelligence Index, second only to Claude Opus 5.5, at Sonnet 5's unchanged $2/$10. Up to high effort it is cheap and fast, answering in seconds where Opus 5.5 takes half a minute. At maximum effort it costs more per task than Opus 5.5 for a lower score.
The short version
Get Claude if
- Teams on Claude Sonnet 5, for whom it is a same-price upgrade
- Chat and interactive products, where it answers in seconds at medium effort
- Everyday coding, bug fixing and documents where high effort is enough
Skip it if
- Anyone planning to run it at max effort: Opus 5.5 reaches the same score for less than half the cost
- Work that needs the top score: Opus 5.5 is ahead at every setting
- Very cheap bulk jobs where a score in the high 30s is enough: GPT-6 Luna costs about a sixth as much
What we tested
- The Artificial Analysis Intelligence Index v4.3.2 table and all five Sonnet 5.5 effort settings, read directly on 29 September 2026
- Anthropic’s launch post, model page, pricing and migration guide
Pros
- Second on the Artificial Analysis Intelligence Index at 56 — behind only Claude Opus 5.5
- Same $2/$10 list price as Sonnet 5 with cache reads at $0.20
- At medium effort beats Sonnet 5’s best score for about a ninth of the cost per task
- At high effort scores 47 for $1.08 a task — a point behind GPT-6 Sol’s best at almost the same cost
- Answers in about 6 seconds end to end at medium where Opus 5.5 needs 18 to 29
- Terminal-Bench 4.0 at 70.6% — ahead of Opus 5.5’s best of 66.4%
- Computer use up from 57.0% to 80.1% on OSWorld 2.1 in Anthropic’s tests
- On every major cloud from launch day with zero data retention available
Cons
- Max effort costs $7.60 a task — more than Opus 5.5 at max for a lower score
- Opus 5.5 at xhigh reaches the same score of 56 for $3.46
- Generated 410M output tokens on the index at max — nearly five times the median
- Opus 5.5 at low scores a point higher than Sonnet 5.5 at medium for less money
- Thinking cannot be disabled — between_tools is the floor and only up to high effort
- Forced tool choice and the old computer-use tool now return errors
- Higher-risk security tasks fall back to Sonnet 5
- Edited conversation histories can invalidate its thinking blocks on newer accounts
Claude Sonnet 5.5 arrived on 28 September at exactly Sonnet 5’s price, $2 in and $10 out per million tokens, with benchmark scores Anthropic put close to Opus 5.5. The independent numbers confirm most of that. On the Artificial Analysis Intelligence Index it scores 56 at maximum effort, second only to Opus 5.5’s 58, and ahead of Claude Fable 5.1 and GPT-6 Astra at 53. Sonnet 5 scored 38 on the same index, so the gain over the model it replaces is 18 points.
Our Opus 5.5 review found that the effort setting decides the verdict more than the model does. That is even more true of Sonnet 5.5, because here the answer changes sides: at some settings it is the best value Anthropic sells, and at others its own sibling beats it on both score and price.
Second on the index, and what each score costs
Artificial Analysis ran Sonnet 5.5 at all five effort settings and publishes what one task cost to complete. It also times the complete response, which it defines as input processing, reasoning and answer generation together. Here is every Sonnet 5.5 and Opus 5.5 setting, ordered by score, with the models they are usually compared against:
| Model and effort | Intelligence | Cost per task | Full response | List price |
|---|---|---|---|---|
| Claude Opus 5.5, max | 58 | $5.98 | 685 s | $4 / $20 |
| Claude Sonnet 5.5, max | 56 | $7.60 | 374 s | $2 / $10 |
| Claude Opus 5.5, xhigh | 56 | $3.46 | 150 s | $4 / $20 |
| Claude Opus 5.5, high | 54 | $1.82 | 62 s | $4 / $20 |
| GPT-6 Astra, max | 53 | $3.26 | 314 s | $10 / $50 |
| Claude Sonnet 5.5, xhigh | 52 | $2.74 | 58 s | $2 / $10 |
| Claude Opus 5.5, medium | 51 | $1.34 | 28.5 s | $4 / $20 |
| GPT-6 Sol, max | 48 | $1.05 | 186 s | $2 / $10 |
| Claude Sonnet 5.5, high (API default) | 47 | $1.08 | 22.6 s | $2 / $10 |
| Claude Opus 5.5, low | 42 | $0.55 | 18.5 s | $4 / $20 |
| Claude Sonnet 5.5, medium (apps default) | 41 | $0.59 | 6.0 s | $2 / $10 |
| Claude Sonnet 5, max | 38 | $5.09 | 190 s | $2 / $10 |
| Claude Sonnet 5.5, low | 36 | $0.41 | 6.8 s | $2 / $10 |
Three readings come out of those numbers.
Up to High, it is excellent value. Medium, the default in Claude’s apps and Claude Code, scores 41 for $0.59 a task. That already beats Sonnet 5’s best score of 38, which cost $5.09, for about a ninth of the price. High, the default on the API, scores 47 for $1.08. That is a point behind GPT-6 Sol at its maximum at almost exactly the same cost, and Sonnet 5.5 has far more headroom above it.
Above High, the value belongs to Opus 5.5. Sonnet 5.5 at Xhigh scores 52 for $2.74, while Opus 5.5 at High scores 54 for $1.82. At Max it reaches 56 for $7.60, and Opus 5.5 reaches the same 56 at Xhigh for $3.46, less than half as much. The reason is token use. At Max, Sonnet 5.5 generated 410 million output tokens across the index, against a median of 88 million and 260 million for Opus 5.5 at its own maximum. A cheaper price per token buys nothing when the model spends more tokens on the same task.
At the bottom of the range, the two overlap. Opus 5.5 at Low scores 42 for $0.55, a point above Sonnet 5.5 at Medium for four cents less. On score and cost alone, Sonnet 5.5 is the cheapest route to a given score only at its Low and High settings. What settles the overlap is the column the index does not score.
Speed is the reason to choose it
Sonnet 5.5 is much faster. At Medium it returns a complete answer in about 6 seconds, with the first output after 1.2 seconds. Opus 5.5 takes 18.5 seconds at Low and 28.5 at Medium. At High the gap is 22.6 seconds against 62.2, and at Max it is about 6 minutes against 11. Output also streams faster at every setting: 85 to 139 tokens a second across Sonnet 5.5’s range, against 74 to 93 for Opus 5.5.
For anything a person waits on, such as a chat reply, a code suggestion or a support answer, that difference outweighs a few cents per task. It is the case for Sonnet 5.5 even at the points where Opus 5.5 is cheaper. Anthropic’s partners report the same pattern in their own tests, though Anthropic chose which to publish. Balyasny Asset Management measured about 121,000 tokens per answer against Sonnet 5’s 497,000 on 2,441 finance tasks, Zendesk says tickets were processed 20% faster, and Box reports 2.4 times the speed with 12% fewer tokens.
Anthropic’s benchmarks
The launch table is Anthropic’s own, run in its own harness. It puts Sonnet 5.5 a few points under Opus 5.5 almost everywhere, with one exception.
That exception is Terminal-Bench 4.0, where Sonnet 5.5 scores 70.6% against Opus 5.5’s best of 66.4%. Anthropic lists Sonnet 5 at 10.3% on the same test and does not explain that figure. On FrontierCode, Anthropic flags a fault of its own: at Max, Sonnet 5.5 more often ran Claude Code’s multi-agent review skill and changed things outside the task, and it scored 46.2% there against 52.1% at Xhigh. Leaving Terminal-Bench aside, the largest gains over Sonnet 5 are in computer use and chart reading, which rise by 23 and 46 points.
Professional work is where Sonnet 5.5 comes closest to Opus 5.5. On GDPval-AA, which covers real tasks from 44 occupations, it finishes two Elo points behind, and on AA-Briefcase eleven. Artificial Analysis ran both boards on a pre-release deployment with a structured-output bug that has since been fixed. Anthropic expects any effect to be small and to understate Sonnet 5.5.
What changes for developers
The price list is identical to Sonnet 5’s: $2.50 per million tokens for five-minute cache writes, $4 for one-hour writes, $0.20 for cache reads, and half price on the Batch API. The tokenizer is also unchanged. The API behaviour is not. Anthropic lists five breaking changes:
- Thinking can be reduced but not disabled.
thinking: {"type": "disabled"}returns a 400 error. The floor is a newbetween_toolsmode, which skips up-front thinking and is accepted only up to High effort. - Forced tool use is gone, as it is on Opus 5.5. Use
autowith strict tool use. - Thinking blocks are tied to the model and the conversation. On accounts created on or after 31 August 2026, replaying one after editing an earlier turn is an error, so keep histories append-only.
- Computer use needs the new toolset on the Claude API and Google Cloud.
- The advisor tool no longer pairs a Sonnet 5.5 executor with Opus 4.8, Opus 4.7 or Sonnet 5.
One further change can make an app go quiet: text written between tool calls now arrives inside thinking blocks. Sonnet 5.5 is also the first Sonnet with the cyber safeguards Anthropic uses on its top models, so higher-risk security tasks visibly fall back to Sonnet 5. For ordinary development none of this shows. For security research it does.
Who should be running it
| If you are… | Do this |
|---|---|
| On Claude Sonnet 5 | Move. Same list price, 18 more index points at max, and at medium it beats Sonnet 5’s best for about a ninth of the cost per task. Budget time for the breaking API changes |
| Building something people wait on | Use it at medium. It returns a complete answer in about 6 seconds, three to five times faster than Opus 5.5 at low or medium |
| Running long agentic jobs | Use high, the API default. If you need more, switch to Opus 5.5 rather than raising the effort: Opus 5.5 at high scores 54 for $1.82, and at xhigh matches Sonnet 5.5’s max score for less than half the cost |
| On Claude Opus 5.5 | Keep it for anything above high effort. Try Sonnet 5.5 where response time matters more than the last few index points |
| On GPT-6 Sol | At high, Sonnet 5.5 is a point behind Sol’s best score at the same cost per task, and answers in 22.6 seconds against 186. Above that, it has headroom Sol does not |
| Running cheap bulk jobs | If a score in the high 30s is enough, GPT-6 Luna at max costs $0.07 a task against $0.41 for Sonnet 5.5 at low |
Verdict
Sonnet 5.5 earns 4.5 out of 5. It is the second-best model on the independent index at the same list price as the model it replaces. It is the fastest way to a strong answer from Claude, and from Low to High effort it is a genuine bargain.
It stops short of Opus 5.5’s 4.9 for one reason, and that reason is Anthropic’s own flagship. Above High, Opus 5.5 reaches every score Sonnet 5.5 can for less money, and at Max the gap is more than double. The breaking API changes are a smaller cost, the same kind Opus 5.5 brought, and worth a day of migration work before switching.
The practical rule fits in one line: Medium for anything a person is waiting on, High for agentic work, and when that is not enough, Opus 5.5 rather than Max. The leaderboard now carries it.
About the score. 4.5/ 5 is BitsMinds' editorial verdict from our own testing and research — not an average of user ratings, which we do not collect. Prices and plan tiers are as published by the vendor on the fact-check date shown above. How we rate →
Related Reviews
Claude Opus 5.5 takes sole first place on the Artificial Analysis Intelligence Index, scoring 58 to the 53 shared by Claude Fable 5.1 and GPT-6 Astra. It lists at $4/$20, a fifth under Opus 5. The better story sits one effort level down: at high it scores 54, above every rival, for $1.82 a task. It also won our three-round build-off, and it was the slowest entrant in every round.
Read review →Anthropic turned the lineup over again: Fable 5.1 replaced Fable 5 at the top on September 1, Opus 5 replaced Opus 4.8 at the same price, and Sonnet 5’s intro pricing has ended. Where Claude leads, where it still struggles, and whether Pro and Max are worth it in September 2026.
Read review →ChatGPT Plus is still the best $20 in AI and Pro is still a $200 question — now with GPT-6 Astra, Sol and Luna rolling out to both tiers. What changed since our GPT-5.5 hands-on, where Claude still wins, and who actually needs Pro.
Read review →Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.