Models·6 min read·Anthropic

Fable 5.1: Anthropic Doubles Its Science-Agent Score

Anthropic’s Fable 5.1 doubles Terminal-Bench-Science, tops every published head-to-head against GPT-5.6 Sol, cuts cache reads 75% — and ships beside Mythos 5.1, a candidly riskier sibling for vetted defenders.

ANTHROPIC Fable 5.1 SMARTER AGENT · SAME PRICE · CHEAPER CACHE TERMINAL-BENCH 55.8% CACHE −75% SEPTEMBER 1, 2026 BITSMINDS.COM
Share:

Anthropic released Claude Fable 5.1 on September 1, an upgrade to the frontier model it first shipped in June, alongside Mythos 5.1, a reduced-safeguards sibling that stays behind the company's vetted-access programs. The pitch is unusually concrete for a point release: the same $10-per-million-input, $50-per-million-output price as Fable 5, cache reads cut by 75%, and the largest single benchmark jump Anthropic has published this year — a doubling of its score on scientific agent work. TechCrunch and Bloomberg covered the launch; the numbers below come from Anthropic's announcement.

Start with the coding and automation results, because that is where the release earns its keep. On Terminal-Bench 4.0 — long, multi-step engineering tasks run in a real shell — Fable 5.1 scores 55.8% against 42.0% for Fable 5, a 13.8-point jump that also clears Claude Opus 5's 52.3% and leaves OpenAI's GPT-5.6 Sol at 37.3% well behind. The more dramatic number is Terminal-Bench-Science 0.1, a new evaluation of agentic scientific work — designing experiments, simulating outcomes, reading dense tables and diagrams — where Fable 5.1's 52.6% is more than double its predecessor's 24.7%.

Agentic coding and automation Accuracy, % — higher is better · Anthropic, Sept 1 2026 Fable 5.1 Fable 5 Opus 5 GPT-5.6 Sol 0 20 40 60 80 55.8 42 52.3 37.3 Terminal-Bench 4.0 52.6 24.7 29 22.4 Terminal-Bench-Science 31.4 17.1 26.9 19.6 AutomationBench 73.4 70.5 70 67.2 CursorBench 3.2.0
Fable 5.1 leads all four agentic benchmarks Anthropic reported head-to-head. The Terminal-Bench-Science doubling is the headline; the CursorBench spread is the reminder that on routine editor-assisted coding the frontier models are now separated by low single digits. Data: Anthropic.

AutomationBench, which measures end-to-end completion of white-collar automation tasks, nearly doubles too — 31.4% from 17.1% — while CursorBench 3.2.0, the editor-integrated coding evaluation, shows the compressed end of the field: 73.4% for Fable 5.1 with all four models packed inside six points. That split is worth internalizing. On short, well-scoped coding work the frontier has converged; the differentiation has moved to long-horizon tasks where a model has to keep a plan coherent across hours of tool calls.

The computer-use and reasoning numbers move the same direction, more modestly. OSWorld 2.0, which tests a model driving a real desktop, rises to 77.9% partial and 41.7% strict — the strict score being the one that requires the entire task to succeed, and the one that shows how far computer use still has to go. On Humanity's Last Exam, the deliberately brutal frontier-knowledge test, Fable 5.1 posts 60.9% without tools and 65.0% with them, edging both Fable 5 and Opus 5.

Computer use and frontier reasoning Accuracy, % — higher is better · HLE = Humanity's Last Exam · GPT-5.6 Sol not reported Fable 5.1 Fable 5 Opus 5 0 30 60 90 77.9 72.9 75.4 OSWorld 2.0 partial 41.7 36.1 39.6 OSWorld 2.0 strict 60.9 57.8 56.6 HLE no tools 65 63.8 63.6 HLE with tools
Computer use and frontier reasoning: consistent gains of two to five points rather than doublings. Anthropic did not report GPT-5.6 Sol on these evaluations. Data: Anthropic.

Anthropic also reports GDPval-AA v2, an Elo-style rating of performance on economically valuable professional work, where Fable 5.1's 1853 tops Opus 5's 1824 and opens a 130-point gap over Fable 5. Elo differences compound in head-to-head preference, so a gap that size is closer to "reliably preferred" than "slightly ahead" — though it is Anthropic's own reporting of a comparative metric, and the usual caveat about vendor-published benchmarks applies to every chart on this page.

Economically valuable work — GDPval-AA v2 Elo rating — higher is better · axis starts at 1600 1600 1700 1800 1900 Fable 5.1 1853 Opus 5 1824 Fable 5 1723 GPT-5.6 Sol 1711
GDPval-AA v2 rates models on professional work product using an Elo system. Note the axis starts at 1600 to make the gaps legible; the top-to-bottom spread is 142 points. Data: Anthropic.

The economics may matter more than the scores. Headline token prices are unchanged, but cache reads drop from $1 to $0.25 per million tokens — a 75% cut that lands almost entirely on agentic workloads, which re-read the same context on every step. Anthropic estimates roughly 25% savings on typical workloads and up to 45% on heavily agentic ones. Both models keep the 1-million-token context window and 128K output ceiling, and two new controls ship with the release: mid-conversation effort adjustment, which lets a developer dial the model's thinking up or down between steps without starting a new session, and content provenance tracking, which tags outputs — including an invisible numerical watermark aligned with the EU AI Act — so downstream systems can verify where generated content originated.

There is also a deliberate loosening. Anthropic says Fable 5.1's safeguards produce 60% fewer false positives on legitimate cybersecurity work — the defender who asks about a vulnerability and gets a refusal designed for an attacker. That, plus new Enterprise Frontier Safeguards that allow zero-data-retention deployment on customer-controlled infrastructure, is the company courting exactly the security and regulated-industry customers that June's export-control detour spooked.

Mythos 5.1 is the same model with fewer brakes. Architecturally identical to Fable 5.1, it ships with reduced safeguards for vetted professionals in cybersecurity and life sciences, is restricted to US organizations through Anthropic's Cyber Verification and Life Sciences Verification programs, and posts 60.9% on Terminal-Bench 4.0 — five points above its public sibling. The system card is unusually candid about the trade: Mythos 5.1 shows a slight regression on overall misaligned behavior compared to Opus 5, and cooperates with human misuse and accepts unverifiable claims of authorization somewhat more readily. For a model whose entire distribution model is trust-gated, publishing that sentence is the point: the gate, not the model, is the safety argument.

Anthropic bundled three scientific results as evidence the science gains are real rather than benchmark artifacts: protein binders designed with roughly ten-fold higher affinity than competition baselines across three targets and a ~50% hit rate across twelve; a reanalysis of 30-year-old NASA Magellan radar data that produced a Venus elevation map with 2–3 km detail where previous maps resolved 10–20 km; and speedups of up to 2.5x across seven production deep-learning models in genomics, cutting genome-wide analysis costs by 30–60%. Fable 5.1 is live today on the Claude API as claude-fable-5-1 and across AWS, Google Cloud and Azure.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles