Claude
Claude
Text & Chatfreemium
4.3

Claude Opus 5 Review: The Smartest Model Anthropic Has Shipped — and the Most Annoying

Opus 5 tops the Artificial Analysis Intelligence Index at 61 and posted a near-4x jump on ARC-AGI-3, all at the same price as Opus 4.8. Then the reviewers who tested it called it 'brilliant but annoying': it argues with instructions, stops early, does unrequested refactors, and breaks prompt scaffolding tuned for Opus 4.8. One more finding deserves attention — a reported jump in its hallucination rate.

Pros

  • Ranked first on the Artificial Analysis Intelligence Index at 61
  • ARC-AGI-3 leap to 30.2% from a 7.8% prior best
  • Within 0.5% of Fable 5 on CursorBench at half the cost
  • Same $5/$25 pricing as the Opus 4.8 it replaces
  • Won a blind 7-model test across six tasks
  • 1M-token context with the freshest knowledge cutoff in the family
  • Best published alignment score of any Anthropic model

Cons

  • Argues with instructions and stops before finishing work
  • Reported hallucination rate up 14 points to 50%
  • Verbose by default — confirmed in Anthropic's own docs
  • Makes unrequested edits and refactors at higher effort
  • Higher effort can reduce coding quality rather than improve it
  • Breaks prompt scaffolding tuned for Opus 4.8
  • Reported to have lost both agentic coding benchmarks to rivals

Claude Opus 5 is the rare model whose reviews split cleanly in two. It sits at the top of the Artificial Analysis Intelligence Index. It also drew a run of write-ups titled "brilliant but annoying" and "hard to love." Both are true, and the gap between them is the most interesting thing about this model.

This review is based on published testing from reviewers who had week-long pre-release access — principally Claire Vo (founder of ChatPRD) and Dan Shipper and Katie Parrott at Every — plus Anthropic's own documentation and system card. Where a finding comes from a single source we say so.

What it is genuinely best at

MeasureResult
Artificial Analysis Intelligence Index61 — first place, ahead of Fable 5's 60
ARC-AGI-3 (novel problem-solving)30.2% versus a 7.8% previous best
CursorBench 3.2Within 0.5% of Fable 5 at roughly half the cost
Blind 7-model test across 6 tasksRanked first by Claire Vo — above Fable 5 and GPT-5.6 Sol
Price$5 / $25 per MTok — unchanged from Opus 4.8

The ARC-AGI-3 result deserves emphasis because it is not incremental. Going from 7.8% to 30.2% on a benchmark built specifically to resist memorisation is the kind of jump that usually only happens once per model generation. And Vo's blind test is meaningful precisely because it was blind — she ranked it first without knowing which model she was scoring.

The complaints are specific and they are real

This is not vague grumbling about vibes. The recurring reports describe the same behaviours:

BehaviourSource
"Argued with instructions and stopped before the work was finished"Dan Shipper — Every
"Didn't play well with our existing skills and plugins"Dan Shipper — Every
Refused to touch a merge conflict during a real coding sessionClaire Vo
Verbosity — "default responses run longer than previous Opus models"Anthropic's own docs
"More changes than the task requires — refactoring or other edits"Anthropic's system card

Two of those five are Anthropic confirming the problem itself, which is what separates this from the usual post-launch grumbling. The model is documented to be more verbose and more prone to doing work you didn't ask for. Reviewers converged on calling the resulting personality "neurotic," and Vo coined "Claude Slop" for the verbosity.

The practical cost is migration friction. Prompt scaffolding tuned for Opus 4.8 does not transfer cleanly. If you have a mature agent harness, budget time to re-tune it rather than swapping the model ID and expecting a free upgrade.

The counterintuitive finding: turn the dial down

The most useful practical discovery came from both testing groups independently: on coding work, lower effort produced better results than higher effort. Reported figures put Opus 5 at 53.4% at medium effort versus 43.6% at xhigh — the inverse of how effort normally scales. Adam Wolff of the Claude Code team is reported to have made medium his default for shipping code.

That matters because it complicates the headline feature. The xhigh setting was one of the launch's marquee additions, and on aggregate intelligence measures raising effort does help. But on agentic coding specifically, giving this model more room appears to let it wander into the unrequested refactors its own system card warns about. Treat the effort dial as something to tune per workload, not to max out.

Caveat: those two figures come from a single secondary write-up rather than a primary benchmark publication, so treat the exact numbers as indicative. The direction is corroborated by two independent testing teams reaching the same "give it less" conclusion.

The number nobody led with

One reported finding deserves more attention than it received: the hallucination rate is said to have risen 14 points to 50%. Also reported — Anthropic initially presented a narrow agentic-coding loss (53.4% against Fable 5's 53.5%) as a win before quietly correcting the chart.

Both come from one source and we could not independently verify either against a primary document, so we are not treating them as settled. But a materially higher hallucination rate on a model this widely deployed is exactly the kind of claim that deserves a direct answer from Anthropic, and Anthropic has not addressed it publicly.

How to actually use it

If you're…Do this
On Opus 4.8 todayUpgrade — same price and better on nearly everything — but re-test your prompts before trusting it in production
Doing agentic codingStart at medium effort. Raise it only if your evals show a gain
Running a tuned agent harnessExpect friction. Budget re-tuning time rather than a drop-in swap
Doing novel reasoning or front-end designThis is where it shines — the ARC-AGI-3 jump is real
Paying for Fable 5 out of habitRe-run your evals. Opus 5 matches it on several measures at half the price
Working where a wrong answer is expensiveVerify outputs until the hallucination question is answered

Verdict

Opus 5 is the most capable model Anthropic has shipped on published aggregate measures, at the same price as the model it replaces. That combination is hard to argue with, and it is why this scores well.

What keeps it from a higher score is that raw capability is not the whole product. A model that argues with instructions, stops early, does unrequested work, and breaks scaffolding built for its predecessor imposes a real cost that no benchmark captures — and Anthropic documents two of those behaviours itself. The reported hallucination increase, if it holds, is a more serious mark still.

The honest summary is the one the early testers landed on: this model is smarter than what came before and more tiring to work with. If your workload rewards raw reasoning, it is the best option available at this price. If your workload depends on precise instruction-following inside an existing pipeline, test carefully before you migrate — and consider staying on a model whose behaviour you have already tuned for.

Related Reviews

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.