Claude Haiku 5.5 Review: Luna’s Price, a Higher Ceiling

Claude Haiku 5.5 scores 43 on the Artificial Analysis Intelligence Index at maximum effort, 26 points above Haiku 4.5, at GPT-6 Luna's $0.10/$0.50. Up to a score of 38, Luna gets there for the same money or less. Above it, Luna has nothing to match Haiku 5.5, which also streams its answers faster.

By BitsMindsHands-on reviewPublished Not re-verified since

The short version

Get Claude if

  • Teams on Claude Haiku 4.5, for whom it is a far better model at a lower cost per task
  • High-volume classification, extraction, summaries and context compaction
  • Subagents under Opus 5.5 or Sonnet 5.5, and computer-use agents on a budget

Skip it if

  • GPT-6 Luna users who are happy with its scores: switching saves nothing per task
  • Products where the first word must appear within a second or two, unless thinking can be switched off
  • Complex agentic coding, where Anthropic itself recommends Sonnet 5.5 or Opus 5.5

What we tested

  • All five Haiku 5.5 effort settings on the Artificial Analysis Intelligence Index v4.3.2, against all five of GPT-6 Luna’s, read directly on 8 October 2026
  • Artificial Analysis’s speed measurements for both, with Claude Haiku 4.5 and Sonnet 5.5 for reference
  • Anthropic’s launch post, model page, pricing page and migration guide
Not re-verified since

Pros

  • Scores 43 on the Artificial Analysis Intelligence Index at max — 26 points above Haiku 4.5
  • Costs less per index task than Haiku 4.5 at every setting while scoring at least 12 points more
  • Reaches 41 and 43 at xhigh and max — scores GPT-6 Luna does not reach at any setting
  • At high effort matches GPT-6 Luna’s best score for about the same cost per task
  • Streams 155 to 244 tokens a second — well ahead of GPT-6 Luna’s 112 to 129
  • List price a tenth of Haiku 4.5’s for prompts up to 100K tokens
  • First Haiku with effort settings and a 1M-token context window
  • OSWorld 2.1 at 72.4% in Anthropic’s tests — up from 15.7% for Haiku 4.5

Cons

  • Up to a score of 38 GPT-6 Luna gets there for the same money or less
  • Took about 13 seconds to a first token at low and medium on its first day of measurements
  • Prompts over 100K tokens cost five times as much per token
  • The same text counts as about 30% more tokens than on Haiku 4.5
  • Max effort burns 440M tokens on the index — more than four times the median
  • Assistant prefill and manual thinking budgets now return errors
  • Non-default temperature or top_p values return an error
  • Anthropic itself steers complex agentic coding to Sonnet 5.5 or Opus 5.5

Claude Haiku 5.5 arrived on 7 October at $0.10 per million input tokens and $0.50 per million output for prompts up to 100,000 tokens. That is a tenth of what Haiku 4.5 costs, and exactly GPT-6 Luna’s price. Anthropic’s own benchmark table put it ahead of Luna on every test the two share. A day later, Artificial Analysis published independent measurements at all five effort settings, and they tell a more nuanced story: a large step past Haiku 4.5 and a higher ceiling than Luna, but not a cheaper route to any score that Luna can already reach.

The independent numbers

On the Artificial Analysis Intelligence Index, version 4.3.2, a set of ten evaluations, Haiku 5.5 scores 43 at maximum effort. Claude Haiku 4.5 scored 17 at its best, so the new model gains 26 points, and it does so for less money per task at every setting. Here is each setting of Haiku 5.5 and GPT-6 Luna, ordered by score, with what one index task cost and how quickly each started and streamed its answer:

Model and effortIntelligenceCost per taskFirst tokenOutput speed
Claude Haiku 5.5, max43$0.21341 s244 tok/s
Claude Haiku 5.5, xhigh41$0.1279 s193 tok/s
Claude Haiku 5.5, high38$0.0827 s179 tok/s
GPT-6 Luna, max38$0.07106 s129 tok/s
GPT-6 Luna, xhigh35$0.0427 s115 tok/s
Claude Haiku 5.5, medium (default)34$0.0513.4 s155 tok/s
GPT-6 Luna, high33$0.0319 s112 tok/s
GPT-6 Luna, medium30$0.02n/an/a
Claude Haiku 5.5, low29$0.0212.8 s185 tok/s
GPT-6 Luna, low22under $0.013.2 s125 tok/s
Claude Haiku 4.5, thinking on17$0.2814.3 s90 tok/s

Costs are as Artificial Analysis publishes them, rounded to the cent. Speed figures are for Anthropic’s and OpenAI’s own APIs with thinking on, and Artificial Analysis had less than a day of Haiku 5.5 measurements when we read them.

Intelligence Index at each effort settingArtificial Analysis Intelligence Index v4.3.2, 0–100, axis from zero · read 8 October 2026Claude Haiku 5.5GPT-6 Luna010203040502922Low3430Medium3833High4135Xhigh4338Max
Haiku 5.5 scores higher than GPT-6 Luna at every setting, by four to seven points. Claude Haiku 4.5 scored 17 at its best on the same index. Data: Artificial Analysis.
What one index task costs at each effort settingUS cents — lower is better · Artificial Analysis, index v4.3.2, rounded as publishedClaude Haiku 5.5GPT-6 Luna05101520252.00.5Low5.02.0Medium8.03.0High12.04.0Xhigh21.07.0Max
Haiku 5.5 costs more than Luna at every setting, and the gap widens at Max, where it generated 440 million tokens across the index to Luna’s 140 million. Luna at Low costs /usr/bin/bash.0045 a task, shown as half a cent. Claude Haiku 4.5 cost 28 cents a task for its 17 points.

Three readings come out of those numbers.

Against Haiku 4.5, there is no contest. The default Medium setting scores twice Haiku 4.5’s best for about a fifth of its cost per task. Even at Max, where Haiku 5.5 generated 440 million tokens across the index, it costs $0.21 a task to Haiku 4.5’s $0.28.

Against Luna, the price is the same and so is the value, up to a point. Haiku 5.5 scores higher than Luna at each matching setting, but each setting also costs more. Compare by score instead and the gap closes: Haiku 5.5 at High scores 38 for $0.08, and Luna at Max scores the same 38 for $0.07. Haiku 5.5 at Medium scores 34 for $0.05, while Luna at Xhigh scores 35 for $0.04. Below a score of 38, Luna is the slightly cheaper buy.

Above 38, Haiku 5.5 has no competition from Luna. Xhigh reaches 41 for $0.12 and Max 43 for $0.21. Luna has no setting that scores above 38. For a team that needs a few more points from a cheap model, that headroom is the case for Haiku 5.5.

Fast once it starts

Anthropic calls Haiku 5.5 its fastest model, and Artificial Analysis bears that out for output. It streamed 155 to 244 tokens a second across the five settings, against 112 to 129 for Luna and about 100 for Sonnet 5.5 at its lower settings.

The wait before the answer begins is another matter. With thinking on, Artificial Analysis measured 12.8 seconds to the first token at Low and 13.4 at Medium, against 3.2 seconds for Luna at Low and about one second for Sonnet 5.5 at Low. These are first-day figures and may settle as more measurements come in. They also include the model’s thinking time. Haiku 5.5, unlike Opus 5.5, still accepts a request with thinking switched off at High effort or below, and Artificial Analysis has not measured that mode. For live customer support, one of the uses Anthropic names, that setting is the one to test first.

Anthropic’s benchmarks

Anthropic’s published benchmarksScore, % — higher is better · self-reported by Anthropic, 7 October 2026Claude Haiku 5.5GPT-6 LunaClaude Haiku 4.502040608010072.448.915.7OSWorld 2.1(offline)39.216.40.0Terminal-Bench4.046.442.4n/rFrontierCode 1.146.429.16.4Chartography
On the tests Anthropic chose, Haiku 5.5 leads Luna by 4 to 24 points. The independent index above, which covers ten evaluations, shows a smaller gap. Data: Anthropic.

On the tests Anthropic published, Haiku 5.5’s lead over Luna is widest in computer use: 72.4% on the offline subset of OSWorld 2.1, against 48.9%. Terminal-Bench 4.0 shows where the small model stops. Haiku 5.5 more than doubles Luna there, but at 39.2% it is far below Sonnet 5.5’s 70.6%, and Anthropic itself recommends its larger models for complex agentic coding. The independent index is the broader measure, and on it the gap to Luna is four to seven points at each setting rather than the double-digit leads in Anthropic’s chart.

What changes for developers

The price depends on prompt length, which no other current Claude model does. Up to 100,000 tokens: $0.10 in, $0.50 out, $0.01 for cache reads. Above that, every rate is five times higher. Haiku 5.5 also uses the newer tokenizer, so the same text counts as about 30% more tokens than on Haiku 4.5. Anthropic lists five breaking changes from Haiku 4.5:

  • Manual extended thinking returns an error. budget_tokens gives way to adaptive thinking and the effort parameter.
  • Non-default sampling parameters return an error. Omit temperature, top_p and top_k.
  • Prefilling the assistant’s reply returns an error. End every request with a user turn.
  • Computer use needs the new toolset on the Claude API and Google Cloud.
  • Editing an earlier turn invalidates thinking blocks, so keep histories append-only.

A response can now open with a thinking block, so code that reads the first content block as the answer needs to select by type. Thinking tokens also count toward max_tokens, which means a tight limit can end a response before any text appears.

Who should be running it

If you are…Do this
On Claude Haiku 4.5Move. At its default Medium setting Haiku 5.5 scores 34 for about $0.05 a task, against 17 for $0.28. Budget a day for the breaking API changes, and recount your prompts: the same text is about 30% more tokens
On GPT-6 LunaStay if Luna’s scores are enough: up to 38, Haiku 5.5 costs the same or a little more per task. Switch if you need more than Luna’s best, or if computer use is the job, where Anthropic’s own tests show the widest lead
Running subagents or compactionStart at High: 38 for about $0.08 a task, and output at 179 tokens a second. Xhigh adds three points for half as much again
Building live chat or supportTest with thinking switched off, which Haiku 5.5 still allows up to High effort. With thinking on, Artificial Analysis measured about 13 seconds to the first token at Low on day one, against about one second for Sonnet 5.5 at Low
Feeding it long documentsWatch the 100K line. Above it every rate is five times higher, though at $0.50 / $2.50 still a quarter of Sonnet 5.5’s
Doing complex agentic codingUse Sonnet 5.5 or Opus 5.5 as the lead and Haiku 5.5 as a helper. It scores 39.2% on Terminal-Bench 4.0 in Anthropic’s tests, against 70.6% for Sonnet 5.5

Verdict

Haiku 5.5 earns 4.3 out of 5. It gains 26 index points over Haiku 4.5 and costs less per task than the model it replaces at every setting. Its best scores are out of GPT-6 Luna’s reach, and it streams its answers faster than Luna once it starts.

It stops short of a higher score for two reasons. Anthropic’s benchmarks promised a clean win over Luna, and the independent index shows a draw on value up to Luna’s best score. And on its first day of measurements, it was slow to begin answering with thinking on, which matters for exactly the live, high-volume work it is sold for. Both may change: speed measurements settle, and thinking can be switched off.

The practical rule: move off Haiku 4.5 now, start at High for agent work, and try thinking off before shipping anything a customer waits on. The leaderboard now carries Haiku 5.5.

About the score. 4.3/ 5 is BitsMinds' editorial verdict from our own testing and research — not an average of user ratings, which we do not collect. Prices and plan tiers are as published by the vendor on the fact-check date shown above. How we rate →

Related Reviews

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.