Fable 5.1: Anthropic Doubles Its Science-Agent Score
Anthropic’s Fable 5.1 doubles Terminal-Bench-Science, tops every published head-to-head against GPT-5.6 Sol, cuts cache reads 75% — and ships beside Mythos 5.1, a candidly riskier sibling for vetted defenders.
Anthropic released Claude Fable 5.1 on September 1, an upgrade to the frontier model it first shipped in June, alongside Mythos 5.1, a reduced-safeguards sibling that stays behind the company's vetted-access programs. The pitch is unusually concrete for a point release: the same $10-per-million-input, $50-per-million-output price as Fable 5, cache reads cut by 75%, and the largest single benchmark jump Anthropic has published this year — a doubling of its score on scientific agent work. TechCrunch and Bloomberg covered the launch; the numbers below come from Anthropic's announcement.
Start with the coding and automation results, because that is where the release earns its keep. On Terminal-Bench 4.0 — long, multi-step engineering tasks run in a real shell — Fable 5.1 scores 55.8% against 42.0% for Fable 5, a 13.8-point jump that also clears Claude Opus 5's 52.3% and leaves OpenAI's GPT-5.6 Sol at 37.3% well behind. The more dramatic number is Terminal-Bench-Science 0.1, a new evaluation of agentic scientific work — designing experiments, simulating outcomes, reading dense tables and diagrams — where Fable 5.1's 52.6% is more than double its predecessor's 24.7%.
AutomationBench, which measures end-to-end completion of white-collar automation tasks, nearly doubles too — 31.4% from 17.1% — while CursorBench 3.2.0, the editor-integrated coding evaluation, shows the compressed end of the field: 73.4% for Fable 5.1 with all four models packed inside six points. That split is worth internalizing. On short, well-scoped coding work the frontier has converged; the differentiation has moved to long-horizon tasks where a model has to keep a plan coherent across hours of tool calls.
The computer-use and reasoning numbers move the same direction, more modestly. OSWorld 2.0, which tests a model driving a real desktop, rises to 77.9% partial and 41.7% strict — the strict score being the one that requires the entire task to succeed, and the one that shows how far computer use still has to go. On Humanity's Last Exam, the deliberately brutal frontier-knowledge test, Fable 5.1 posts 60.9% without tools and 65.0% with them, edging both Fable 5 and Opus 5.
Anthropic also reports GDPval-AA v2, an Elo-style rating of performance on economically valuable professional work, where Fable 5.1's 1853 tops Opus 5's 1824 and opens a 130-point gap over Fable 5. Elo differences compound in head-to-head preference, so a gap that size is closer to "reliably preferred" than "slightly ahead" — though it is Anthropic's own reporting of a comparative metric, and the usual caveat about vendor-published benchmarks applies to every chart on this page.
The economics may matter more than the scores. Headline token prices are unchanged, but cache reads drop from $1 to $0.25 per million tokens — a 75% cut that lands almost entirely on agentic workloads, which re-read the same context on every step. Anthropic estimates roughly 25% savings on typical workloads and up to 45% on heavily agentic ones. Both models keep the 1-million-token context window and 128K output ceiling, and two new controls ship with the release: mid-conversation effort adjustment, which lets a developer dial the model's thinking up or down between steps without starting a new session, and content provenance tracking, which tags outputs — including an invisible numerical watermark aligned with the EU AI Act — so downstream systems can verify where generated content originated.
There is also a deliberate loosening. Anthropic says Fable 5.1's safeguards produce 60% fewer false positives on legitimate cybersecurity work — the defender who asks about a vulnerability and gets a refusal designed for an attacker. That, plus new Enterprise Frontier Safeguards that allow zero-data-retention deployment on customer-controlled infrastructure, is the company courting exactly the security and regulated-industry customers that June's export-control detour spooked.
Mythos 5.1 is the same model with fewer brakes. Architecturally identical to Fable 5.1, it ships with reduced safeguards for vetted professionals in cybersecurity and life sciences, is restricted to US organizations through Anthropic's Cyber Verification and Life Sciences Verification programs, and posts 60.9% on Terminal-Bench 4.0 — five points above its public sibling. The system card is unusually candid about the trade: Mythos 5.1 shows a slight regression on overall misaligned behavior compared to Opus 5, and cooperates with human misuse and accepts unverifiable claims of authorization somewhat more readily. For a model whose entire distribution model is trust-gated, publishing that sentence is the point: the gate, not the model, is the safety argument.
Anthropic bundled three scientific results as evidence the science gains are real rather than benchmark artifacts: protein binders designed with roughly ten-fold higher affinity than competition baselines across three targets and a ~50% hit rate across twelve; a reanalysis of 30-year-old NASA Magellan radar data that produced a Venus elevation map with 2–3 km detail where previous maps resolved 10–20 km; and speedups of up to 2.5x across seven production deep-learning models in genomics, cutting genome-wide analysis costs by 30–60%. Fable 5.1 is live today on the Claude API as claude-fable-5-1 and across AWS, Google Cloud and Azure.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.