Part of our AI Comparisons
Models·7 min read·BitsMinds Analysis

Claude Opus 5 vs Fable 5 vs Sonnet 5 vs Opus 4.8: The Benchmark Comparison

Opus 5 costs the same as the Opus 4.8 it replaces, beats the double-priced Fable 5 on Anthropic's own coding charts, and burns a fraction of the tokens doing it. We lined up every published number — specs, prices, benchmarks — and flagged exactly where the comparisons don't hold.

BITSMINDS ANALYSIS · CLAUDE LINEUP Opus 5 vs Fable 5 vs Sonnet 5 What actually changed — on price, specs, and the benchmarks Claude Fable 5 $10 / $50 33.7% Claude Opus 5 $5 / $25 43.3% Claude Sonnet 5 $3 / $15 not measured Opus 4.8 · legacy $5 / $25 18.7% Bars: Frontier-Bench v0.1 (coding), as published by Anthropic All three current models share a 1M-token context · prices per million input / output tokens
Share:

A BitsMinds analysis. Claude Opus 5 landed on July 24, and it arrived into an unusually crowded family: it replaces Opus 4.8, sits below Fable 5 in Anthropic's own hierarchy, and costs more than Sonnet 5. So which one should you actually be calling? We pulled every published number and laid them side by side.

How we compared them — and what's missing

Two caveats matter more here than in a typical shootout, and they shape everything below.

First, Anthropic published its Opus 5 charts as images, not tables. The launch post states results in relative terms ("more than doubles Opus 4.8's performance," "within 0.5% of Fable 5's peak") rather than as a numeric grid. Where exact figures exist below, they come from Anthropic's own charts as transcribed by third parties, from the ARC Prize Foundation, or from public leaderboards — each labeled by source.

Second, these four models were never benchmarked against each other on one shared suite. Sonnet 5 launched June 30 against Terminal-Bench 2.1, SWE-bench Pro and OSWorld-Verified. Opus 5 launched July 24 against Frontier-Bench, CursorBench 3.2 and OSWorld 2.0 — different evals, different versions. The one model present in both sets is Opus 4.8, so we use it as the bridge: the honest way to compare Opus 5 and Sonnet 5 today is through how each one measures against the model they both replace or sit beside.

The spec sheet

Start with what isn't in dispute. These figures come straight from Anthropic's model documentation.

SpecClaude Opus 5Claude Fable 5Claude Sonnet 5Opus 4.8 (legacy)
API model IDclaude-opus-5claude-fable-5claude-sonnet-5claude-opus-4-8
Price / MTok (in / out)$5 / $25$10 / $50$3 / $15$5 / $25
Context window1M tokens1M tokens1M tokens1M tokens
Max output128k (300k via Batch beta)128k128k128k
Comparative latencyModerateSlowerFastModerate
Effort levelsAll five (incl. xhigh)All fiveAll fiveAll five
Reliable knowledge cutoffMay 2026Jan 2026Jan 2026Jan 2026
StatusCurrent · default on MaxCurrent · top tierCurrentLegacy

Two things jump out. Every current model now shares the same 1M-token context — capacity has stopped being a differentiator across the lineup. And Opus 5 has the freshest knowledge cutoff in the family by four months, including a fresher one than the nominally more capable Fable 5. If your work touches anything that happened in early 2026, that gap is more practically useful than a benchmark point.

One pricing footnote: Sonnet 5's $3/$15 is its standard rate, but introductory pricing of $2/$10 runs through August 31, 2026.

Opus 5 vs Opus 4.8: the straight upgrade

This is the cleanest comparison of the four, because the price is identical — $5/$25 either way — so every gain is free. It is also the most lopsided.

Benchmark / measureClaude Opus 5Claude Opus 4.8Source
Frontier-Bench v0.1 (coding)43.3%18.7%Anthropic
ARC-AGI 3 (novel problem-solving)30.2%1.5%ARC Prize Foundation
SWE-bench Verified~97%88.6%vals.ai leaderboard
Organic chemistry+10.2 ptsbaselineAnthropic
Protein tasks+7.7 ptsbaselineAnthropic
Reasoning tokens (trading eval)~1/7 as manybaselineAnthropic
Latency (same eval)under halfbaselineAnthropic
Token use, legal work at max effort26% fewerbaselineAnthropic
Misaligned-behavior score (lower is better)2.30higherAnthropic audit
Price / MTok$5 / $25$5 / $25Anthropic

The efficiency numbers deserve as much attention as the capability ones. Anthropic reports Opus 5 using roughly a seventh of the reasoning tokens at under half the latency of Opus 4.8 on a trading evaluation, and 26% fewer tokens on legal work at max effort. Since you pay per token, a model that is both better and more economical at identical list pricing makes staying on Opus 4.8 difficult to justify. Anthropic's own alignment audit also scores Opus 5 lower on misaligned behavior than Opus 4.8, Sonnet 5 and Fable 5 — the best of the family on that measure.

One caveat on the SWE-bench Verified figure: independent trackers listed Opus 5 between 96% and 97% within a day of launch, and at least one leaderboard carried a "last updated" date that predates the launch itself. Treat it as directionally right, not settled.

Opus 5 vs Fable 5: near-frontier at half the price

Here's the awkward one. Anthropic still describes Fable 5 as its "most capable widely released model" and prices it at double. But Opus 5 beats it on several published measures.

Benchmark / measureClaude Opus 5Claude Fable 5
Frontier-Bench v0.1 (coding)43.3%33.7%
SWE-bench Verified~97%95.0%
CursorBench 3.2 (Opus 5 at max effort)within 0.5% of peak, half the cost per taskpeak score
OSWorld 2.0 (computer use)beats Fable 5's best, ~1/3 the costprior best
Price / MTok$5 / $25$10 / $50
Comparative latencyModerateSlower
Reliable knowledge cutoffMay 2026Jan 2026
Anthropic's positioning"complex agentic coding and enterprise work""most capable widely released model"

Read that last row against the rest and the tension is obvious. On the coding and computer-use evals Anthropic chose to publish, the cheaper model won or drew — while the docs still route you to Fable 5 for "the highest available capability." Both can be true: Fable 5 was built for long-running agentic work and retains headroom on tasks that don't appear on these charts, and Anthropic notably did not publish a Fable 5 number for ARC-AGI 3, where Opus 5 posted its most striking result.

The practical read: Fable 5 is no longer the automatic choice just because it costs more. For coding and computer-use work specifically, Opus 5 is the better-value default, and Fable 5 becomes the escalation path you reach for when your own evals show it earns the premium.

Opus 5 vs Sonnet 5: no shared benchmark, so mind the bridge

These two have never been run against each other on a common eval. What we can do is line each up against Opus 4.8, which both were measured against, and read the gap.

BenchmarkClaude Sonnet 5Opus 4.8 (shared baseline)Claude Opus 5
Terminal-Bench 2.180.474.6not published
SWE-bench Pro63.269.2not published
Humanity's Last Exam (with tools)57.457.9not published
OSWorld-Verified81.283.4measured on OSWorld 2.0 instead
GDPval-AA v21,6181,615"state of the art," no figure given
Frontier-Bench v0.1not measured18.7%43.3%
Price / MTok$3 / $15$5 / $25$5 / $25
Comparative latencyFastModerateModerate

Sonnet 5 traded roughly evenly with Opus 4.8 — ahead on Terminal-Bench and GDPval, behind on SWE-bench Pro and OSWorld — at 60% of the price and noticeably faster. Opus 5 then more than doubled Opus 4.8 on Frontier-Bench. Chain those together and Opus 5 is clearly the stronger model, but Sonnet 5 remains the value and latency play, and the gap on general reasoning is far narrower than the gap on hard agentic coding.

The verdict: which one should you use?

If you're…UseWhy
Still on Opus 4.8Switch to Opus 5Same price, better on every published measure, and materially cheaper per task in tokens.
Doing agentic coding or computer useOpus 5 (try xhigh)Its strongest published category, at half Fable 5's rate.
Paying for Fable 5 out of habitRe-run your evalsOpus 5 matched or beat it on the coding and computer-use charts Anthropic published.
On genuinely frontier, long-horizon agent workKeep Fable 5Still Anthropic's designated top tier, with headroom these charts don't capture.
Running high volume or latency-sensitive workSonnet 5Fast tier, $3/$15 standard — and $2/$10 through Aug 31, 2026.
Building chat, classification, or subagentsSonnet 5 at low/medium effortEffort is the cost dial; frontier capability is wasted here.

For how Anthropic's lineup stacks up against the rest of the field, see our cross-lab benchmark comparison.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

ANTHROPIC · OFFICIAL LAUNCH Claude Opus 5 Is Here The rumors were right — right down to the “xhigh” mode 1M-token context new “xhigh” effort mode same price as Opus 4.8 Launched July 24, 2026 · replaces Opus 4.8 · Fable 5 remains the top-tier model
Models

Claude Opus 5 Is Officially Here — 1M Context, a New 'xhigh' Mode, and the Rumors Were Right

GOOGLE · GEMINI Three Flash models ship… …while Gemini 3.5 Pro is still nowhere to be seen 3.6 Flash −17% output tokens 3.5 Flash-Lite cheap · beats Gemini 3 3.5 Flash Cyber finds & fixes bugs Gemini 3.5 Pro promised June · still not shipped → Meanwhile, Google says it has begun its biggest pretraining run yet — Gemini 4 Shipped July 21 via the Gemini API & Gemini Enterprise · the price war grinds on
Models

Google Ships Three New Gemini Flash Models — but 3.5 Pro Is Still Missing, and It's Already Pretraining Gemini 4

RUMOR WATCH · UNCONFIRMED Claude Opus 5 — the evidence board Cursor · Jul 9 “Claude Honeycomb EAP” — then gone Vertex AI · ×2 listing spotted twice, vanished both times 5? the next flagship? 1M context? leaked, unverified “xhigh” mode? launch “this week”?? No official API ID · no system card · every thread on this board is still a rumor
Models

Claude Opus 5 Rumor Watch: The 'Honeycomb' Leak, a Reported 1M Context, and Talk of a Launch This Week