Claude Sonnet 5.5 Is Out: Sonnet Price, Near-Opus Scores
Anthropic released Claude Sonnet 5.5 on 28 September at Sonnet 5's price of $2 and $10 per million tokens, with benchmark scores close to Opus 5.5. Independent tests back the pitch up to high effort and show where it stops: at maximum effort it costs more per task than Opus 5.5.
Claude Sonnet 5.5 is Anthropic's new mid-tier model, released on 28 September 2026 as the second member of the Claude 5.5 family after Opus 5.5. It answers to the model ID claude-sonnet-5-5, has a 1M-token context window with up to 128K tokens of output, and costs $2 per million input tokens and $10 per million output, exactly what Claude Sonnet 5 costs. It is available now on the Claude API, in Claude's apps and Claude Code, and through Amazon Bedrock, Google Cloud and Microsoft Foundry. Haiku 5.5, the third model in the family, is still due “in the coming weeks”.
Anthropic's pitch is a clear upgrade on Sonnet 5 that generates output more than 30% faster and costs up to 30% less per task, because it needs far fewer tokens to do the same work. Where Opus 5.5 is aimed at complex work that needs careful judgement, Sonnet 5.5 is positioned for well-scoped everyday tasks, bug fixing, and polished documents, slides and spreadsheets.
Close to Opus, far past Sonnet 5
The next three charts use Anthropic's own figures from its launch post. Independent measurements follow further down.
Coding is where the jump is largest. On Terminal-Bench 4.0, Anthropic puts Sonnet 5.5 at 70.6%, ahead of Opus 5.5's best of 66.4%. It lists Sonnet 5 at 10.3% on the same test and does not comment on that figure, which is far below the 52.3% it published for Opus 5 on 22 September. On CursorBench, built from real Cursor sessions, Sonnet 5.5's best score is within about two points of Opus 5.5. FrontierCode, which asks whether a code change could be merged without human edits, has a wrinkle Anthropic flags itself: Sonnet 5.5 scores 52.1% at Xhigh effort but 46.2% at Max, because at Max it more often ran Claude Code's multi-agent review skill and made changes beyond the task.
Away from code, Sonnet 5.5 lands 1.7 to 3.2 points under Opus 5.5 on Humanity's Last Exam, OSWorld 2.1 and Chartography. The movement against Sonnet 5 is the bigger story: computer use rises from 57.0% to 80.1%, and chart recognition from 15.6% to 61.6%. Anthropic also says it is the first Sonnet model to beat Pokémon Red working only from screenshots.
The Elo boards come from Artificial Analysis, which Anthropic says ran them on a pre-release deployment with a since-fixed bug affecting structured outputs; it expects any effect to be small and to understate Sonnet 5.5. On GDPval-AA, which covers tasks from 44 occupations, Sonnet 5.5 finishes two points behind Opus 5.5 and about 400 ahead of Sonnet 5.
The independent numbers: cheap up to High
Artificial Analysis has already run Sonnet 5.5 through its Intelligence Index, version 4.3.2, a set of ten evaluations, at all five effort settings. At maximum effort it scores 56, second only to Opus 5.5 among the models on the index: two points behind it, and three ahead of Fable 5.1 and GPT-6 Astra. The same run cost $7.60 per index task, more than Opus 5.5's $5.98 at its own maximum. The curve below that point tells a more useful story than the headline.
From Low to High, Sonnet 5.5 is cheap: $0.41, $0.59 and $1.08 a task. Medium, the default in Claude's apps and Claude Code, scores 41, above Sonnet 5's best of 38 at Max for about a ninth of Sonnet 5's $5.09. That independently supports Anthropic's claim that Sonnet 5.5 at low or medium effort beats Sonnet 5's best for around a tenth of the cost. High, the default on the API, scores 47 for $1.08, a point behind GPT-6 Sol at its maximum (48 for $1.05) at almost the same price.
Above High, the value drains away. Xhigh costs $2.74 a task for 52, and Max $7.60 for 56. Opus 5.5 reaches 54 at High for $1.82 and 56 at Xhigh for $3.46, less than half what Sonnet 5.5 pays for the same score. The two models overlap at the bottom of the range as well: Opus 5.5 at Low scores 42 for $0.55, a point more than Sonnet 5.5 at Medium for four cents less. On this index Sonnet 5.5 is the cheapest route to a given score only at its Low and High settings; everywhere else Opus 5.5 gets there for less. One index is not every workload, and a job's own token use can differ, but the result argues against reaching for Max by default.
What the index does not price is speed. Artificial Analysis also times the complete response, reasoning included. At Medium, Sonnet 5.5 returns in about 6 seconds, against 18.5 seconds for Opus 5.5 at Low and 28.5 at Medium. Its output streams faster at every setting too, 139 tokens a second at Max against 93 for Opus 5.5. For interactive work, where waiting costs more than tokens, that is the case for Sonnet 5.5 even at the points where the index says Opus 5.5 is cheaper.
What changes for developers
The price list is identical to Sonnet 5's: $2.50 per million tokens for five-minute cache writes, $4 for one-hour writes, $0.20 for cache reads, and half price on the Batch API, where output can run to 300K tokens with a beta header. The tokenizer is also unchanged, so the same text counts the same tokens. Behaviour is what changes. Anthropic lists five breaking changes for code written for Sonnet 5:
- Thinking can be reduced but not disabled. Adaptive thinking is on by default, and
thinking: {"type": "disabled"}now returns a 400 error. The lowest setting is a newbetween_toolsmode, which turns off up-front thinking but still returns short progress notes between tool calls. It is accepted only at High effort or below. - Forced tool use is gone. A
tool_choiceofanyor a named tool returns an error, as it already does on Opus 5.5. Anthropic points to strict tool use or structured outputs instead. - Thinking blocks are tied to the model and the conversation. No other model can read Sonnet 5.5's thinking blocks, and on accounts created on or after 31 August 2026, replaying one after an earlier turn has been edited returns an error. Anthropic's advice is to keep conversations append-only.
- Computer use moves to the new toolset,
computer_toolset_20260801, on the Claude API and Google Cloud. - The advisor tool drops three pairings. Opus 4.8, Opus 4.7 and Sonnet 5 can no longer advise a Sonnet 5.5 executor.
A quieter change can catch streaming interfaces: text Sonnet 5.5 writes between tool calls now arrives inside thinking blocks, so an app that shows it to users goes silent until it opts back in. Non-default temperature, top_p or top_k values return an error, and effort defaults to High on the API and Medium in Claude's apps and Claude Code. The migration guide has before-and-after requests for each change.
The first Sonnet with cyber safeguards
Anthropic describes Sonnet 5.5's cyber capabilities as comparable to Opus 5's, and ships it with the kind of cybersecurity safeguards it has so far reserved for its most capable models, a first for a Sonnet. Routine bug-finding and fixing is unaffected, but higher-risk security tasks visibly fall back to Sonnet 5. It is also the first Sonnet to launch with classifiers that block reasoning extraction, aimed at distillation. In its launch post, Anthropic says Sonnet 5.5 matches or improves on Sonnet 5 on most alignment measures in an automated audit of about 1,850 scenarios, and is the least likely of any of its models to probe the limits of its container; the evaluations are detailed in its system card. That claim arrives two days after OpenAI said it had paused work on its most capable models when an agent escaped its sandbox.
How the leaks held up
Our preview, published on the morning of 28 September, weighed a noisy leak chain against the one source with a hit behind it: the X account @lyraxana, which had called Opus 5.5's name, date and prices. Lyra said Sonnet 5.5 would come that week, possibly before Monday morning was out. It shipped that Monday. The preview noted that Monday morning had passed without it, but that line went up at 11:11 UTC, a little after 4 a.m. in San Francisco, and the model arrived later the same day.
| Claim | Where it came from | Outcome |
|---|---|---|
| Release in the week of 28 September | @lyraxana on X | Right: 28 September |
| Launch on 30 September or 1 October | X account Pankaj Kumar, via Startup Fortune | Wrong: two days early |
| $2 input, $10 output, $0.20 cache reads | TestingCatalog | Right |
| 872K-token context, 128K output | Factory's Droid registry | Output right; the context window is 1M |
| 2M-token context | Earlier X posts | Wrong |
| Adaptive thinking on by default | TestingCatalog | Right, though unlike Opus 5.5 it can be turned down to between_tools |
| Codename “Fennec” | X posts | Still unconfirmed |
The price was the claim the preview said would look the same whether the leak was genuine or copied from Sonnet 5's price sheet, and it held: Sonnet 5.5 costs exactly what Sonnet 5 does. The 872K figure, which the preview flagged as a number no current Anthropic model uses, did not become one.
Taken together, launch day supports a narrower conclusion than either the announcement or the index alone. Sonnet 5.5 is a real step past Sonnet 5 at the same list price, and at Medium or High effort it is one of the cheapest ways to a strong score on independent tests. Above High, on the numbers available so far, the better buy is not a higher effort setting but Opus 5.5.
Our verdict: Sonnet 5.5 earns 4.5 out of 5. Read the Claude Sonnet 5.5 review, which sets out when to choose it over Opus 5.5.
More on Claude
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.