Claude Review 2026: Plans, Models and Verdict
Anthropic turned the lineup over again: Fable 5.1 replaced Fable 5 at the top on September 1, Opus 5 replaced Opus 4.8 at the same price, and Sonnet 5’s intro pricing has ended. Where Claude leads, where it still struggles, and whether Pro and Max are worth it in September 2026.
Pros
- Best-in-class nuanced writing
- Fable 5.1 leads long-horizon agentic coding (55.8% Terminal-Bench 4.0)
- Opus 5 tops the Artificial Analysis index at Opus 4.8’s price
- Sonnet 5 is a capable default at $3/$15
- 1M-token context with reliable recall
- Claude Code Desktop is a productivity multiplier
- "Most honest" — flags its own uncertainty
Cons
- Fable 5.1 is the most expensive frontier model per task ($10/$50 list and heavy token use)
- Opus 5 argues with instructions and needs prompt re-tuning
- Hard safety guardrails still block some legitimate cyber/bio work
- Still no native image generation
- Top Mythos-class capability gated behind vetted access
The Bottom Line
Claude remains our team's primary AI for serious work — and Anthropic spent the last two months widening its lead. Since our May review the lineup has turned over twice: Opus 4.8 and then Opus 5 at the flagship-workhorse tier, Fable 5 and then Fable 5.1 at the frontier, and Sonnet 5 as the default that makes everyday Claude both smarter and cheaper. Claude still leads on the three things that matter most — writing quality, code-review depth, and intellectual honesty — and the premium is easier to justify now that Sonnet 5 handles most work at a fraction of the price.
Current Lineup (September 2026)
- Claude Fable 5.1 (September 1, 2026) — The frontier model, an upgrade to Fable 5 at the same $10 / $50 per million tokens, with cache reads cut 75% (from $1 to $0.25). The headline jump is long-horizon agentic work: 55.8% on Terminal-Bench 4.0 against 42.0% for Fable 5, 52.3% for Opus 5 and 37.3% for GPT-5.6 Sol, and a doubling on Terminal-Bench-Science to 52.6%. It edges Opus 5 on Humanity's Last Exam (65.0% with tools) and on GDPval-AA v2 (1853 vs 1824). Same 1M context and 128K output, plus mid-conversation effort adjustment. Anthropic says its safeguards produce 60% fewer false-positive refusals on legitimate cybersecurity work. Full numbers in our launch coverage.
- Claude Opus 5 (July 2026) — Replaced Opus 4.8 at the same $5 / $25. It topped the Artificial Analysis Intelligence Index at 61 on launch, ahead of Fable 5's 60 (index version of July 2026; on the current v4.3 index Opus 5 scores 51 and Fable 5 50 — see our leaderboard), and nearly quadrupled the previous best on ARC-AGI-3 (30.2% vs 7.8%). It also earned a reputation: reviewers called it brilliant but annoying — it argues with instructions and wanders into unrequested refactors, prompt scaffolding tuned for 4.8 does not transfer cleanly, and the Claude Code team's own default for shipping code is medium effort, not max. We scored it 4.3 in our Opus 5 review.
- Claude Sonnet 5 (June 30, 2026) — The mid-tier and the default on Free and Pro. The most agentic Sonnet yet (63.2% on SWE-Bench Pro), matching Opus 4.8 on some knowledge work. Its introductory $2/$10 pricing ended on August 31; it is now $3 / $15.
- Claude Haiku 4.5 — Still the fast, cheap option at $1 / $5 for high-volume, latency-sensitive tasks.
- Claude Mythos 5 and 5.1 — The frontier models Fable is built from. Not public. Mythos 5.1, architecturally identical to Fable 5.1 with reduced safeguards, is restricted to vetted US organizations in cybersecurity and life sciences.
What Claude Does Best
Writing — Still Unmatched
For longform content, marketing copy, technical writing, or anything requiring tonal nuance, Claude consistently produces output that needs less editing than competitors. The model has a recognizable voice — slightly playful, willing to push back, comfortable with ambiguity. We test the same prompts across Claude, GPT-5.6, and Gemini regularly. For "writing that doesn't sound AI," Claude still wins ~70% of the time.
Code Review and Refactoring
Claude's models dominate code-review benchmarks — Fable 5 led public models on SWE-Bench Pro at 80.3, Fable 5.1 extends the lead on long-horizon work with 55.8% on Terminal-Bench 4.0 against 52.3% for Opus 5 and 37.3% for GPT-5.6 Sol, and even mid-tier Sonnet 5 clears 63.2% on SWE-Bench Pro. In real use, Claude catches subtle bugs, suggests architectural improvements, and explains trade-offs in a way that feels like working with a senior engineer. Dynamic Workflows, introduced with Opus 4.8, fan a large review out across hundreds of parallel subagents.
Long Document Analysis
The 1M-token context window is now table stakes (Gemini matches it), but Claude's recall across long contexts remains the most reliable. Drop in a 200-page contract and ask "find every clause that limits liability and rank by risk" — you get a real answer with citations.
Vision
High-resolution vision carried over and improved through Opus 4.8 and Fable 5: feed in UI screenshots, dense charts, and document scans and you get accurate analysis rather than a vague summary. Fable 5.1 pushes computer use further still — 77.9% on OSWorld 2.0 (partial) against 72.9% for Fable 5 and 75.4% for Opus 5.
Claude Code Desktop App — Still the Differentiator
Anthropic's Claude Code Desktop app remains the single biggest workflow advantage in the lineup. A multi-session sidebar lets you run several concurrent agentic tasks across different projects; drag-and-drop layout, an integrated terminal, an in-app file editor, and a rebuilt diff viewer make it feel like a real IDE for AI coding rather than a chat tool retrofitted to do code. With Opus 5 underneath, a single instruction can spawn and coordinate a swarm of subagents — start it at medium effort for shipping code, which is the Claude Code team's own default, and save max for genuinely hard problems.
For pure agentic work — "describe a feature, get it built across files" — Claude Code is best-in-class. Cursor still wins on inline autocomplete, but for multi-file, multi-step builds Claude leads our coding workflow.
Where It Falls Short
Pricing
The top of the range is expensive: Fable 5.1 is $10 / $50 per million tokens, unchanged from Fable 5 although cache reads fell 75%. It shares that list price with GPT-6 Astra but uses far more tokens per task — Artificial Analysis puts the cost of running its Intelligence Index at $10,816 for Fable 5.1 against $4,064 for Astra — so the premium is real. Opus 5 sits in the middle at $5 / $25, and Sonnet 5 is $3 / $15 now that the introductory $2/$10 has ended. The winning strategy is unchanged: default to Sonnet 5 for most work and reserve Fable 5.1 or Opus 5 for the hard problems where the quality gap actually pays for itself.
Strict Content Policies
Claude refuses some legitimate requests other models accept, and Fable 5 went further — its classifier guardrails blocked cyber, bio, and model-distillation requests outright, quietly falling back to Opus 4.8. Anthropic says Fable 5.1 produces 60% fewer of those false positives on legitimate cybersecurity work, and routes vetted professionals to the reduced-safeguards Mythos 5.1 instead. For security research, content-moderation work, and creative writing involving conflict, you'll still occasionally hit walls. More context usually unblocks it, but it's friction.
No Image Generation
If you need image creation, you'll still pair Claude with ChatGPT, Midjourney, or FLUX. Anthropic remains uninterested in entering the image-generation space.
Plans
- Free — Daily Sonnet 5 usage. Good for evaluation.
- Pro ($20/mo) — Sonnet 5 as the default plus access to the flagship tier, where Opus 5 and Fable 5.1 have succeeded the Opus 4.8 and Fable 5 we tested, along with Projects, large context, and file uploads.
- Max ($100 / $200/mo) — 5× or 20× Pro limits, heavier flagship use, and full Claude Code access.
- Team / Enterprise — Org admin, SSO, audit logs, and custom data residency.
Plan prices are as published by Anthropic. Per-plan model allocation shifts with every release, so check the plan page for the current split before choosing a tier on the strength of one model.
Claude vs GPT-6 Astra vs Gemini
Claude wins: writing quality, deep code review, careful reasoning, and agentic coding via Claude Code. On the independent scorecard Fable 5.1 led Astra 66 to 61 on the Artificial Analysis Intelligence Index as published on September 3, 2026 (on the re-based v4.3 index both score 53), and on Humanity's Last Exam with tools it scores 65.0% to Astra's 57.2%.
GPT-6 Astra wins: computer use and cyber — a perfect score on ExploitBench and OSWorld tasks in about 40 minutes where GPT-5.6 Sol needed 75 — plus raw token efficiency, which is how it matches Fable 5's coding score at less than half the cost despite the same $10/$50 list price. By OpenAI's own self-reported numbers it also edges Fable 5.1 by 1.9 points on Terminal-Bench 4.0. Our Astra review weighs those claims.
Gemini wins: Google Workspace integration, native multimodal (video and audio), and free-tier value.
Verdict
Claude is still the best AI for serious text and code work in September 2026. What's changed since May is the value equation: Sonnet 5 makes the Pro plan cheaper to lean on for everyday work, while Fable 5.1 delivers frontier-class capability when you need it and Opus 5 covers most of the ground at half the price. Pro at $20/month remains essential for anyone who writes or codes professionally; Max is justified for power users running Claude Code and the flagship models hard. The premium at the top is real — Fable 5.1 is the most expensive frontier model per task — but the writing, code review, and honesty still lead the field. Score: 4.8/5.
Further Reading: Claude Guides
- Claude: The Complete Guide to Anthropic's AI — the full walkthrough of models, plans, and Claude Code.
- How to Use Claude Fable 5: Access, Costs, and When to Choose It — when the premium model is worth it.
- What Is Claude Cowork? Anthropic's Agentic Workspace — how Claude moves beyond chat to read, edit, and act.
- Claude Code: A Guide to the Terminal-Native AI Coding Agent — the coding agent this review raves about.
- How to Use Claude Code Routines — schedule an AI agent to run itself in the cloud.
About the score. 4.8/ 5 is BitsMinds' editorial verdict from our own testing and research — not an average of user ratings, which we do not collect. Prices and plan tiers are as published by the vendor on the fact-check date shown above. How we rate →
Related Reviews
ChatGPT Plus is still the best $20 in AI and Pro is still a $200 question — now with GPT-6 Astra rolling out to both tiers. What changed since our GPT-5.5 hands-on, where Claude still wins, and who actually needs Pro.
Read review →Anthropic's Claude Fable 5 is the most capable model the public can use today — topping SWE-Bench Pro and excelling at vision and long tasks. It is also the priciest major model and ships with hard safety guardrails. Our first-look verdict.
Read review →Perplexity combines real-time web search with AI synthesis and proper citations. After daily use for two years, here's why it's our team's research default and where it still falls short.
Read review →Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.