Claude Opus 5 vs Fable 5 vs Sonnet 5 vs Opus 4.8: The Benchmark Comparison
Opus 5 costs the same as the Opus 4.8 it replaces, beats the double-priced Fable 5 on Anthropic's own coding charts, and burns a fraction of the tokens doing it. We lined up every published number — specs, prices, benchmarks — and flagged exactly where the comparisons don't hold.
A BitsMinds analysis. Claude Opus 5 landed on July 24, and it arrived into an unusually crowded family: it replaces Opus 4.8, sits below Fable 5 in Anthropic's own hierarchy, and costs more than Sonnet 5. So which one should you actually be calling? We pulled every published number and laid them side by side.
How we compared them — and what's missing
Two caveats matter more here than in a typical shootout, and they shape everything below.
First, Anthropic published its Opus 5 charts as images, not tables. The launch post states results in relative terms ("more than doubles Opus 4.8's performance," "within 0.5% of Fable 5's peak") rather than as a numeric grid. Where exact figures exist below, they come from Anthropic's own charts as transcribed by third parties, from the ARC Prize Foundation, or from public leaderboards — each labeled by source.
Second, these four models were never benchmarked against each other on one shared suite. Sonnet 5 launched June 30 against Terminal-Bench 2.1, SWE-bench Pro and OSWorld-Verified. Opus 5 launched July 24 against Frontier-Bench, CursorBench 3.2 and OSWorld 2.0 — different evals, different versions. The one model present in both sets is Opus 4.8, so we use it as the bridge: the honest way to compare Opus 5 and Sonnet 5 today is through how each one measures against the model they both replace or sit beside.
The spec sheet
Start with what isn't in dispute. These figures come straight from Anthropic's model documentation.
| Spec | Claude Opus 5 | Claude Fable 5 | Claude Sonnet 5 | Opus 4.8 (legacy) |
|---|---|---|---|---|
| API model ID | claude-opus-5 | claude-fable-5 | claude-sonnet-5 | claude-opus-4-8 |
| Price / MTok (in / out) | $5 / $25 | $10 / $50 | $3 / $15 | $5 / $25 |
| Context window | 1M tokens | 1M tokens | 1M tokens | 1M tokens |
| Max output | 128k (300k via Batch beta) | 128k | 128k | 128k |
| Comparative latency | Moderate | Slower | Fast | Moderate |
| Effort levels | All five (incl. xhigh) | All five | All five | All five |
| Reliable knowledge cutoff | May 2026 | Jan 2026 | Jan 2026 | Jan 2026 |
| Status | Current · default on Max | Current · top tier | Current | Legacy |
Two things jump out. Every current model now shares the same 1M-token context — capacity has stopped being a differentiator across the lineup. And Opus 5 has the freshest knowledge cutoff in the family by four months, including a fresher one than the nominally more capable Fable 5. If your work touches anything that happened in early 2026, that gap is more practically useful than a benchmark point.
One pricing footnote: Sonnet 5's $3/$15 is its standard rate, but introductory pricing of $2/$10 runs through August 31, 2026.
Opus 5 vs Opus 4.8: the straight upgrade
This is the cleanest comparison of the four, because the price is identical — $5/$25 either way — so every gain is free. It is also the most lopsided.
| Benchmark / measure | Claude Opus 5 | Claude Opus 4.8 | Source |
|---|---|---|---|
| Frontier-Bench v0.1 (coding) | 43.3% | 18.7% | Anthropic |
| ARC-AGI 3 (novel problem-solving) | 30.2% | 1.5% | ARC Prize Foundation |
| SWE-bench Verified | ~97% | 88.6% | vals.ai leaderboard |
| Organic chemistry | +10.2 pts | baseline | Anthropic |
| Protein tasks | +7.7 pts | baseline | Anthropic |
| Reasoning tokens (trading eval) | ~1/7 as many | baseline | Anthropic |
| Latency (same eval) | under half | baseline | Anthropic |
| Token use, legal work at max effort | 26% fewer | baseline | Anthropic |
| Misaligned-behavior score (lower is better) | 2.30 | higher | Anthropic audit |
| Price / MTok | $5 / $25 | $5 / $25 | Anthropic |
The efficiency numbers deserve as much attention as the capability ones. Anthropic reports Opus 5 using roughly a seventh of the reasoning tokens at under half the latency of Opus 4.8 on a trading evaluation, and 26% fewer tokens on legal work at max effort. Since you pay per token, a model that is both better and more economical at identical list pricing makes staying on Opus 4.8 difficult to justify. Anthropic's own alignment audit also scores Opus 5 lower on misaligned behavior than Opus 4.8, Sonnet 5 and Fable 5 — the best of the family on that measure.
One caveat on the SWE-bench Verified figure: independent trackers listed Opus 5 between 96% and 97% within a day of launch, and at least one leaderboard carried a "last updated" date that predates the launch itself. Treat it as directionally right, not settled.
Opus 5 vs Fable 5: near-frontier at half the price
Here's the awkward one. Anthropic still describes Fable 5 as its "most capable widely released model" and prices it at double. But Opus 5 beats it on several published measures.
| Benchmark / measure | Claude Opus 5 | Claude Fable 5 |
|---|---|---|
| Frontier-Bench v0.1 (coding) | 43.3% | 33.7% |
| SWE-bench Verified | ~97% | 95.0% |
| CursorBench 3.2 (Opus 5 at max effort) | within 0.5% of peak, half the cost per task | peak score |
| OSWorld 2.0 (computer use) | beats Fable 5's best, ~1/3 the cost | prior best |
| Price / MTok | $5 / $25 | $10 / $50 |
| Comparative latency | Moderate | Slower |
| Reliable knowledge cutoff | May 2026 | Jan 2026 |
| Anthropic's positioning | "complex agentic coding and enterprise work" | "most capable widely released model" |
Read that last row against the rest and the tension is obvious. On the coding and computer-use evals Anthropic chose to publish, the cheaper model won or drew — while the docs still route you to Fable 5 for "the highest available capability." Both can be true: Fable 5 was built for long-running agentic work and retains headroom on tasks that don't appear on these charts, and Anthropic notably did not publish a Fable 5 number for ARC-AGI 3, where Opus 5 posted its most striking result.
The practical read: Fable 5 is no longer the automatic choice just because it costs more. For coding and computer-use work specifically, Opus 5 is the better-value default, and Fable 5 becomes the escalation path you reach for when your own evals show it earns the premium.
Opus 5 vs Sonnet 5: no shared benchmark, so mind the bridge
These two have never been run against each other on a common eval. What we can do is line each up against Opus 4.8, which both were measured against, and read the gap.
| Benchmark | Claude Sonnet 5 | Opus 4.8 (shared baseline) | Claude Opus 5 |
|---|---|---|---|
| Terminal-Bench 2.1 | 80.4 | 74.6 | not published |
| SWE-bench Pro | 63.2 | 69.2 | not published |
| Humanity's Last Exam (with tools) | 57.4 | 57.9 | not published |
| OSWorld-Verified | 81.2 | 83.4 | measured on OSWorld 2.0 instead |
| GDPval-AA v2 | 1,618 | 1,615 | "state of the art," no figure given |
| Frontier-Bench v0.1 | not measured | 18.7% | 43.3% |
| Price / MTok | $3 / $15 | $5 / $25 | $5 / $25 |
| Comparative latency | Fast | Moderate | Moderate |
Sonnet 5 traded roughly evenly with Opus 4.8 — ahead on Terminal-Bench and GDPval, behind on SWE-bench Pro and OSWorld — at 60% of the price and noticeably faster. Opus 5 then more than doubled Opus 4.8 on Frontier-Bench. Chain those together and Opus 5 is clearly the stronger model, but Sonnet 5 remains the value and latency play, and the gap on general reasoning is far narrower than the gap on hard agentic coding.
The verdict: which one should you use?
| If you're… | Use | Why |
|---|---|---|
| Still on Opus 4.8 | Switch to Opus 5 | Same price, better on every published measure, and materially cheaper per task in tokens. |
| Doing agentic coding or computer use | Opus 5 (try xhigh) | Its strongest published category, at half Fable 5's rate. |
| Paying for Fable 5 out of habit | Re-run your evals | Opus 5 matched or beat it on the coding and computer-use charts Anthropic published. |
| On genuinely frontier, long-horizon agent work | Keep Fable 5 | Still Anthropic's designated top tier, with headroom these charts don't capture. |
| Running high volume or latency-sensitive work | Sonnet 5 | Fast tier, $3/$15 standard — and $2/$10 through Aug 31, 2026. |
| Building chat, classification, or subagents | Sonnet 5 at low/medium effort | Effort is the cost dial; frontier capability is wasted here. |
For how Anthropic's lineup stacks up against the rest of the field, see our cross-lab benchmark comparison.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.