Claude Haiku 5.5 Review: Luna’s Price, a Higher Ceiling
Claude Haiku 5.5 scores 43 on the Artificial Analysis Intelligence Index at maximum effort, 26 points above Haiku 4.5, at GPT-6 Luna's $0.10/$0.50. Up to a score of 38, Luna gets there for the same money or less. Above it, Luna has nothing to match Haiku 5.5, which also streams its answers faster.
The short version
Get Claude if
- Teams on Claude Haiku 4.5, for whom it is a far better model at a lower cost per task
- High-volume classification, extraction, summaries and context compaction
- Subagents under Opus 5.5 or Sonnet 5.5, and computer-use agents on a budget
Skip it if
- GPT-6 Luna users who are happy with its scores: switching saves nothing per task
- Products where the first word must appear within a second or two, unless thinking can be switched off
- Complex agentic coding, where Anthropic itself recommends Sonnet 5.5 or Opus 5.5
What we tested
- All five Haiku 5.5 effort settings on the Artificial Analysis Intelligence Index v4.3.2, against all five of GPT-6 Luna’s, read directly on 8 October 2026
- Artificial Analysis’s speed measurements for both, with Claude Haiku 4.5 and Sonnet 5.5 for reference
- Anthropic’s launch post, model page, pricing page and migration guide
Pros
- Scores 43 on the Artificial Analysis Intelligence Index at max — 26 points above Haiku 4.5
- Costs less per index task than Haiku 4.5 at every setting while scoring at least 12 points more
- Reaches 41 and 43 at xhigh and max — scores GPT-6 Luna does not reach at any setting
- At high effort matches GPT-6 Luna’s best score for about the same cost per task
- Streams 155 to 244 tokens a second — well ahead of GPT-6 Luna’s 112 to 129
- List price a tenth of Haiku 4.5’s for prompts up to 100K tokens
- First Haiku with effort settings and a 1M-token context window
- OSWorld 2.1 at 72.4% in Anthropic’s tests — up from 15.7% for Haiku 4.5
Cons
- Up to a score of 38 GPT-6 Luna gets there for the same money or less
- Took about 13 seconds to a first token at low and medium on its first day of measurements
- Prompts over 100K tokens cost five times as much per token
- The same text counts as about 30% more tokens than on Haiku 4.5
- Max effort burns 440M tokens on the index — more than four times the median
- Assistant prefill and manual thinking budgets now return errors
- Non-default temperature or top_p values return an error
- Anthropic itself steers complex agentic coding to Sonnet 5.5 or Opus 5.5
Claude Haiku 5.5 arrived on 7 October at $0.10 per million input tokens and $0.50 per million output for prompts up to 100,000 tokens. That is a tenth of what Haiku 4.5 costs, and exactly GPT-6 Luna’s price. Anthropic’s own benchmark table put it ahead of Luna on every test the two share. A day later, Artificial Analysis published independent measurements at all five effort settings, and they tell a more nuanced story: a large step past Haiku 4.5 and a higher ceiling than Luna, but not a cheaper route to any score that Luna can already reach.
The independent numbers
On the Artificial Analysis Intelligence Index, version 4.3.2, a set of ten evaluations, Haiku 5.5 scores 43 at maximum effort. Claude Haiku 4.5 scored 17 at its best, so the new model gains 26 points, and it does so for less money per task at every setting. Here is each setting of Haiku 5.5 and GPT-6 Luna, ordered by score, with what one index task cost and how quickly each started and streamed its answer:
| Model and effort | Intelligence | Cost per task | First token | Output speed |
|---|---|---|---|---|
| Claude Haiku 5.5, max | 43 | $0.21 | 341 s | 244 tok/s |
| Claude Haiku 5.5, xhigh | 41 | $0.12 | 79 s | 193 tok/s |
| Claude Haiku 5.5, high | 38 | $0.08 | 27 s | 179 tok/s |
| GPT-6 Luna, max | 38 | $0.07 | 106 s | 129 tok/s |
| GPT-6 Luna, xhigh | 35 | $0.04 | 27 s | 115 tok/s |
| Claude Haiku 5.5, medium (default) | 34 | $0.05 | 13.4 s | 155 tok/s |
| GPT-6 Luna, high | 33 | $0.03 | 19 s | 112 tok/s |
| GPT-6 Luna, medium | 30 | $0.02 | n/a | n/a |
| Claude Haiku 5.5, low | 29 | $0.02 | 12.8 s | 185 tok/s |
| GPT-6 Luna, low | 22 | under $0.01 | 3.2 s | 125 tok/s |
| Claude Haiku 4.5, thinking on | 17 | $0.28 | 14.3 s | 90 tok/s |
Costs are as Artificial Analysis publishes them, rounded to the cent. Speed figures are for Anthropic’s and OpenAI’s own APIs with thinking on, and Artificial Analysis had less than a day of Haiku 5.5 measurements when we read them.
Three readings come out of those numbers.
Against Haiku 4.5, there is no contest. The default Medium setting scores twice Haiku 4.5’s best for about a fifth of its cost per task. Even at Max, where Haiku 5.5 generated 440 million tokens across the index, it costs $0.21 a task to Haiku 4.5’s $0.28.
Against Luna, the price is the same and so is the value, up to a point. Haiku 5.5 scores higher than Luna at each matching setting, but each setting also costs more. Compare by score instead and the gap closes: Haiku 5.5 at High scores 38 for $0.08, and Luna at Max scores the same 38 for $0.07. Haiku 5.5 at Medium scores 34 for $0.05, while Luna at Xhigh scores 35 for $0.04. Below a score of 38, Luna is the slightly cheaper buy.
Above 38, Haiku 5.5 has no competition from Luna. Xhigh reaches 41 for $0.12 and Max 43 for $0.21. Luna has no setting that scores above 38. For a team that needs a few more points from a cheap model, that headroom is the case for Haiku 5.5.
Fast once it starts
Anthropic calls Haiku 5.5 its fastest model, and Artificial Analysis bears that out for output. It streamed 155 to 244 tokens a second across the five settings, against 112 to 129 for Luna and about 100 for Sonnet 5.5 at its lower settings.
The wait before the answer begins is another matter. With thinking on, Artificial Analysis measured 12.8 seconds to the first token at Low and 13.4 at Medium, against 3.2 seconds for Luna at Low and about one second for Sonnet 5.5 at Low. These are first-day figures and may settle as more measurements come in. They also include the model’s thinking time. Haiku 5.5, unlike Opus 5.5, still accepts a request with thinking switched off at High effort or below, and Artificial Analysis has not measured that mode. For live customer support, one of the uses Anthropic names, that setting is the one to test first.
Anthropic’s benchmarks
On the tests Anthropic published, Haiku 5.5’s lead over Luna is widest in computer use: 72.4% on the offline subset of OSWorld 2.1, against 48.9%. Terminal-Bench 4.0 shows where the small model stops. Haiku 5.5 more than doubles Luna there, but at 39.2% it is far below Sonnet 5.5’s 70.6%, and Anthropic itself recommends its larger models for complex agentic coding. The independent index is the broader measure, and on it the gap to Luna is four to seven points at each setting rather than the double-digit leads in Anthropic’s chart.
What changes for developers
The price depends on prompt length, which no other current Claude model does. Up to 100,000 tokens: $0.10 in, $0.50 out, $0.01 for cache reads. Above that, every rate is five times higher. Haiku 5.5 also uses the newer tokenizer, so the same text counts as about 30% more tokens than on Haiku 4.5. Anthropic lists five breaking changes from Haiku 4.5:
- Manual extended thinking returns an error.
budget_tokensgives way to adaptive thinking and the effort parameter. - Non-default sampling parameters return an error. Omit
temperature,top_pandtop_k. - Prefilling the assistant’s reply returns an error. End every request with a user turn.
- Computer use needs the new toolset on the Claude API and Google Cloud.
- Editing an earlier turn invalidates thinking blocks, so keep histories append-only.
A response can now open with a thinking block, so code that reads the first content block as the answer needs to select by type. Thinking tokens also count toward max_tokens, which means a tight limit can end a response before any text appears.
Who should be running it
| If you are… | Do this |
|---|---|
| On Claude Haiku 4.5 | Move. At its default Medium setting Haiku 5.5 scores 34 for about $0.05 a task, against 17 for $0.28. Budget a day for the breaking API changes, and recount your prompts: the same text is about 30% more tokens |
| On GPT-6 Luna | Stay if Luna’s scores are enough: up to 38, Haiku 5.5 costs the same or a little more per task. Switch if you need more than Luna’s best, or if computer use is the job, where Anthropic’s own tests show the widest lead |
| Running subagents or compaction | Start at High: 38 for about $0.08 a task, and output at 179 tokens a second. Xhigh adds three points for half as much again |
| Building live chat or support | Test with thinking switched off, which Haiku 5.5 still allows up to High effort. With thinking on, Artificial Analysis measured about 13 seconds to the first token at Low on day one, against about one second for Sonnet 5.5 at Low |
| Feeding it long documents | Watch the 100K line. Above it every rate is five times higher, though at $0.50 / $2.50 still a quarter of Sonnet 5.5’s |
| Doing complex agentic coding | Use Sonnet 5.5 or Opus 5.5 as the lead and Haiku 5.5 as a helper. It scores 39.2% on Terminal-Bench 4.0 in Anthropic’s tests, against 70.6% for Sonnet 5.5 |
Verdict
Haiku 5.5 earns 4.3 out of 5. It gains 26 index points over Haiku 4.5 and costs less per task than the model it replaces at every setting. Its best scores are out of GPT-6 Luna’s reach, and it streams its answers faster than Luna once it starts.
It stops short of a higher score for two reasons. Anthropic’s benchmarks promised a clean win over Luna, and the independent index shows a draw on value up to Luna’s best score. And on its first day of measurements, it was slow to begin answering with thinking on, which matters for exactly the live, high-volume work it is sold for. Both may change: speed measurements settle, and thinking can be switched off.
The practical rule: move off Haiku 4.5 now, start at High for agent work, and try thinking off before shipping anything a customer waits on. The leaderboard now carries Haiku 5.5.
About the score. 4.3/ 5 is BitsMinds' editorial verdict from our own testing and research — not an average of user ratings, which we do not collect. Prices and plan tiers are as published by the vendor on the fact-check date shown above. How we rate →
Related Reviews
Claude Opus 5.5 Review: The New No. 1 Costs Less
Claude Opus 5.5 takes sole first place on the Artificial Analysis Intelligence Index, scoring 58 to the 53 shared by Claude Fable 5.1 and GPT-6 Astra. It lists at $4/$20, a fifth under Opus 5. The better story sits one effort level down: at high it scores 54, above every rival, for $1.82 a task. It also won our three-round build-off, and it was the slowest entrant in every round.
Claude Review 2026: Plans, Models and Verdict
Anthropic turned the lineup over again: Fable 5.1 replaced Fable 5 at the top on September 1, Opus 5 replaced Opus 4.8 at the same price, and Sonnet 5’s intro pricing has ended. Where Claude leads, where it still struggles, and whether Pro and Max are worth it in September 2026.
ChatGPT Plus & Pro Review 2026: Is It Worth $20?
ChatGPT Plus is still the best $20 in AI and Pro is still a $200 question — now with GPT-6 Astra, Sol and Luna rolling out to both tiers. What changed since our GPT-5.5 hands-on, where Claude still wins, and who actually needs Pro.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.