Claude Opus 5 Review: The Smartest Model Anthropic Has Shipped — and the Most Annoying
Opus 5 tops the Artificial Analysis Intelligence Index at 61 and posted a near-4x jump on ARC-AGI-3, all at the same price as Opus 4.8. Then the reviewers who tested it called it 'brilliant but annoying': it argues with instructions, stops early, does unrequested refactors, and breaks prompt scaffolding tuned for Opus 4.8. One more finding deserves attention — a reported jump in its hallucination rate.
Pros
- Ranked first on the Artificial Analysis Intelligence Index at 61
- ARC-AGI-3 leap to 30.2% from a 7.8% prior best
- Within 0.5% of Fable 5 on CursorBench at half the cost
- Same $5/$25 pricing as the Opus 4.8 it replaces
- Won a blind 7-model test across six tasks
- 1M-token context with the freshest knowledge cutoff in the family
- Best published alignment score of any Anthropic model
Cons
- Argues with instructions and stops before finishing work
- Reported hallucination rate up 14 points to 50%
- Verbose by default — confirmed in Anthropic's own docs
- Makes unrequested edits and refactors at higher effort
- Higher effort can reduce coding quality rather than improve it
- Breaks prompt scaffolding tuned for Opus 4.8
- Reported to have lost both agentic coding benchmarks to rivals
Claude Opus 5 is the rare model whose reviews split cleanly in two. It sits at the top of the Artificial Analysis Intelligence Index. It also drew a run of write-ups titled "brilliant but annoying" and "hard to love." Both are true, and the gap between them is the most interesting thing about this model.
This review is based on published testing from reviewers who had week-long pre-release access — principally Claire Vo (founder of ChatPRD) and Dan Shipper and Katie Parrott at Every — plus Anthropic's own documentation and system card. Where a finding comes from a single source we say so.
What it is genuinely best at
| Measure | Result |
|---|---|
| Artificial Analysis Intelligence Index | 61 — first place, ahead of Fable 5's 60 |
| ARC-AGI-3 (novel problem-solving) | 30.2% versus a 7.8% previous best |
| CursorBench 3.2 | Within 0.5% of Fable 5 at roughly half the cost |
| Blind 7-model test across 6 tasks | Ranked first by Claire Vo — above Fable 5 and GPT-5.6 Sol |
| Price | $5 / $25 per MTok — unchanged from Opus 4.8 |
The ARC-AGI-3 result deserves emphasis because it is not incremental. Going from 7.8% to 30.2% on a benchmark built specifically to resist memorisation is the kind of jump that usually only happens once per model generation. And Vo's blind test is meaningful precisely because it was blind — she ranked it first without knowing which model she was scoring.
The complaints are specific and they are real
This is not vague grumbling about vibes. The recurring reports describe the same behaviours:
| Behaviour | Source |
|---|---|
| "Argued with instructions and stopped before the work was finished" | Dan Shipper — Every |
| "Didn't play well with our existing skills and plugins" | Dan Shipper — Every |
| Refused to touch a merge conflict during a real coding session | Claire Vo |
| Verbosity — "default responses run longer than previous Opus models" | Anthropic's own docs |
| "More changes than the task requires — refactoring or other edits" | Anthropic's system card |
Two of those five are Anthropic confirming the problem itself, which is what separates this from the usual post-launch grumbling. The model is documented to be more verbose and more prone to doing work you didn't ask for. Reviewers converged on calling the resulting personality "neurotic," and Vo coined "Claude Slop" for the verbosity.
The practical cost is migration friction. Prompt scaffolding tuned for Opus 4.8 does not transfer cleanly. If you have a mature agent harness, budget time to re-tune it rather than swapping the model ID and expecting a free upgrade.
The counterintuitive finding: turn the dial down
The most useful practical discovery came from both testing groups independently: on coding work, lower effort produced better results than higher effort. Reported figures put Opus 5 at 53.4% at medium effort versus 43.6% at xhigh — the inverse of how effort normally scales. Adam Wolff of the Claude Code team is reported to have made medium his default for shipping code.
That matters because it complicates the headline feature. The xhigh setting was one of the launch's marquee additions, and on aggregate intelligence measures raising effort does help. But on agentic coding specifically, giving this model more room appears to let it wander into the unrequested refactors its own system card warns about. Treat the effort dial as something to tune per workload, not to max out.
Caveat: those two figures come from a single secondary write-up rather than a primary benchmark publication, so treat the exact numbers as indicative. The direction is corroborated by two independent testing teams reaching the same "give it less" conclusion.
The number nobody led with
One reported finding deserves more attention than it received: the hallucination rate is said to have risen 14 points to 50%. Also reported — Anthropic initially presented a narrow agentic-coding loss (53.4% against Fable 5's 53.5%) as a win before quietly correcting the chart.
Both come from one source and we could not independently verify either against a primary document, so we are not treating them as settled. But a materially higher hallucination rate on a model this widely deployed is exactly the kind of claim that deserves a direct answer from Anthropic, and Anthropic has not addressed it publicly.
How to actually use it
| If you're… | Do this |
|---|---|
| On Opus 4.8 today | Upgrade — same price and better on nearly everything — but re-test your prompts before trusting it in production |
| Doing agentic coding | Start at medium effort. Raise it only if your evals show a gain |
| Running a tuned agent harness | Expect friction. Budget re-tuning time rather than a drop-in swap |
| Doing novel reasoning or front-end design | This is where it shines — the ARC-AGI-3 jump is real |
| Paying for Fable 5 out of habit | Re-run your evals. Opus 5 matches it on several measures at half the price |
| Working where a wrong answer is expensive | Verify outputs until the hallucination question is answered |
Verdict
Opus 5 is the most capable model Anthropic has shipped on published aggregate measures, at the same price as the model it replaces. That combination is hard to argue with, and it is why this scores well.
What keeps it from a higher score is that raw capability is not the whole product. A model that argues with instructions, stops early, does unrequested work, and breaks scaffolding built for its predecessor imposes a real cost that no benchmark captures — and Anthropic documents two of those behaviours itself. The reported hallucination increase, if it holds, is a more serious mark still.
The honest summary is the one the early testers landed on: this model is smarter than what came before and more tiring to work with. If your workload rewards raw reasoning, it is the best option available at this price. If your workload depends on precise instruction-following inside an existing pipeline, test carefully before you migrate — and consider staying on a model whose behaviour you have already tuned for.
Related Reviews
Anthropic's lineup leveled up: Fable 5 is the most capable public model, Opus 4.8 runs hundreds of parallel subagents, and Sonnet 5 is the new Pro default at a fraction of the cost. Where Claude leads, where it still struggles, and whether Pro/Max are worth it in mid-2026.
Read review →GPT-5.5 launched April 23, 2026 — OpenAI's most capable and intuitive model yet. After three weeks of intensive use, here's where GPT-5.5 dominates, where Claude still wins, and whether ChatGPT Pro at $200/month is justified.
Read review →Anthropic's Claude Fable 5 is the most capable model the public can use today — topping SWE-Bench Pro and excelling at vision and long tasks. It is also the priciest major model and ships with hard safety guardrails. Our first-look verdict.
Read review →Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.