Claude Haiku 5.5 Is Out at a Tenth of Haiku 4.5's Price
Anthropic released Claude Haiku 5.5 on 7 October at $0.10 and $0.50 per million tokens for prompts up to 100,000 tokens, a tenth of Haiku 4.5's price. That is exactly what GPT-6 Luna costs, and on Anthropic's own benchmarks Haiku 5.5 beats Luna on every test the two share.
Claude Haiku 5.5 is Anthropic's new small model, released on 7 October 2026 the third Claude 5.5 model, after Opus 5.5 on 22 September and Sonnet 5.5 on 28 September. Its model ID is claude-haiku-5-5. It has a 1M-token context window and up to 128K tokens of output, up from 200K and 64K on Haiku 4.5, and costs $0.10 per million input tokens and $0.50 per million output for prompts up to 100,000 tokens. It is available now on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS.
Anthropic calls it the cheapest, fastest and most capable small model it has released. The jobs it names are high-volume and repetitive: summaries, context compaction, database queries and classification, plus work as a subagent under Opus 5.5 or Sonnet 5.5 on coding tasks. Speed is the other half of the pitch. Haiku 5.5 is Anthropic's fastest model at standard speed, though a footnote concedes that Opus models in Fast Mode are quicker, and Anthropic points to live customer support and browser automation as the work where that speed pays.
GPT-6 Luna's price, to the cent
The price lands exactly on OpenAI's. GPT-6 Luna, OpenAI's budget model since 22 September, also costs $0.10 per million input tokens and $0.50 per million output, with cached input at $0.01, and Haiku 5.5's cache reads cost the same $0.01 for prompts under the 100K mark. Anthropic put Luna in its own comparison, and on its own table Haiku 5.5 is ahead on every test where both have a score.
The next three charts are Anthropic's figures, run and selected by the company that sells the model. No independent numbers exist yet: Artificial Analysis had not published Haiku 5.5 when this was written. On its Intelligence Index, version 4.3.2, GPT-6 Luna at maximum effort scores 38 and costs $0.07 per index task, the bar Haiku 5.5 will have to clear on cost per task rather than on price per token.
Computer use is the widest gap. On the offline subset of OSWorld 2.1, where an agent operates a real computer through long multi-step tasks, Haiku 5.5 scores 72.4% to Luna's 48.9% and Haiku 4.5's 15.7%, within about 12 points of Sonnet 5.5. Terminal-Bench 4.0 shows the limit. Haiku 5.5 more than doubles Luna there at 39.2%, but Sonnet 5.5 reaches 70.6%, and Anthropic says plainly that Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding. FrontierCode, which asks whether a change could be merged without human edits, is closer: 46.4% against Luna's 42.4%.
The reasoning and vision rows show how far a single generation moved the small model. Haiku 4.5 scored 6.4% on Chartography; Haiku 5.5 scores 46.4%, against Luna's 29.1%. On Humanity's Last Exam, a test of expert-level academic knowledge, it gets 45.9% without tools and 57.4% with them.
The Elo boards come from Artificial Analysis and rate professional work; GDPval-AA, for instance, draws its tasks from 44 occupations. Haiku 5.5 scores 1620 on it, 183 points ahead of Luna and 220 behind Sonnet 5.5.
Where the 75% comes from
Anthropic's headline is that Haiku 5.5 costs about 75% less to run than Haiku 4.5. The per-token cut is larger and the per-task figure is smaller, for two reasons spelled out in the footnotes. First, the price depends on prompt length. Up to 100,000 tokens, every rate is a tenth of Haiku 4.5's. Above that, input costs $0.50 and output $2.50, still half of Haiku 4.5's $1 and $5, but five times Haiku 5.5's own short-prompt rate. Anthropic says 90% of requests to Haiku 4.5 fell under the line. Second, Haiku 5.5 uses the newer tokenizer from Sonnet 5.5 and Opus 5.5. The announcement says this means slightly more tokens per task; the developer docs put it at about 30% more tokens for the same text than Haiku 4.5.
| Per million tokens | Input | Cache read | Output |
|---|---|---|---|
| Claude Haiku 5.5, prompts up to 100K | $0.10 | $0.01 | $0.50 |
| Claude Haiku 5.5, prompts over 100K | $0.50 | $0.05 | $2.50 |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 |
| Claude Sonnet 5.5 | $2.00 | $0.10 | $10.00 |
The tiered price is new in the current lineup. Anthropic's pricing page says every model from Claude 4.6 on bills a 900K-token request at the same rate as a 9K one, with Haiku 5.5 named as the exception. Cache writes follow the same split: $0.125 for five minutes and $0.20 for an hour under the line, $0.625 and $1 above it. The Batch API halves both tiers and accepts up to 300K output tokens with a beta header. For classification, routing and short summaries, the 100K line will rarely matter. For retrieval jobs that stuff long documents into each call, it decides the bill.
Effort settings reach Haiku
Haiku 5.5 is the first Haiku with an adjustable effort setting, from Low to Max, so developers can trade answer quality against speed and cost as they can on the larger models. It defaults to Medium. Adaptive thinking is on by default, but unlike Opus 5.5, Haiku 5.5 still lets a request turn thinking off, at High effort or below. Anthropic's launch post plots each effort setting against cost per attempt on OSWorld, GDPval-AA and Humanity's Last Exam, but it does not print the values behind the dots.
Moving from Haiku 4.5 is not a model-name swap. Anthropic's list of changes includes five that break existing code:
- Manual extended thinking returns an error.
budget_tokensis gone; adaptive thinking and the effort parameter replace it. - Non-default sampling parameters return an error. Requests have to omit
temperature,top_pandtop_k. - Prefilling the assistant's reply returns an error. Every request has to end with a user turn.
- Computer use needs the new toolset,
computer_toolset_20260801in place ofcomputer_20250124, on the Claude API and Google Cloud. - Editing an earlier turn invalidates thinking blocks. Conversations that send thinking blocks back have to stay append-only.
Quieter changes can still trip a working app. A response can now begin with a thinking block, so code that reads the first content block as the answer has to select blocks by type. Thinking text is omitted unless a request asks for a summary. Safety classifiers can decline a request with a refusal stop reason, with no server-side fallback. Thinking tokens count toward max_tokens, so a small limit can end a response before any text appears. The migration guide covers each change.
What early customers reported
Anthropic published results from six companies that tested the model before launch, each on its own workload. HubSpot says Haiku 5.5 posted the best score it has seen on its CRM evaluation suite, 92.8% averaged over three runs. AlphaSense ran 400 queries through Ask in Document, a feature it says handles about 8 million calls a week, and measured 0.84 against Haiku 4.5's 0.76. Box reports a score 11 points above Haiku 4.5 at about half the latency, and Asana a cut of more than 30% in task-completion latency against the model it uses today. Cognition says its Devin Fusion setup, with Opus 5.5 leading and Haiku 5.5 as the helper, holds a FrontierCode score of 66.2 at lower cost and latency. These are the customers Anthropic chose to quote, on tests they designed.
Safeguards: between Haiku 4.5 and Sonnet 5.5
Anthropic reports large improvements over Haiku 4.5 on almost all of its alignment evaluations, with far fewer cases of misaligned behaviour and less willingness to help with misuse; the system card has the detail. Its cybersecurity safeguards are stricter than Haiku 4.5's but looser than those on Anthropic's other recent models: they allow a wider range of defensive work than Sonnet 5.5's, while still blocking penetration testing and other techniques more often used by attackers. Biology safeguards match Sonnet 5, Sonnet 5.5 and Opus 5.
Also announced: cheaper Sonnet 5.5 caching and API credit for subscribers
Anthropic paired the launch with two price moves. Cache reads on Sonnet 5.5 are halved from $0.20 to $0.10 per million tokens from 7 October, which Anthropic estimates makes Sonnet 5.5 about 20% cheaper on most agentic work, because cached context makes up a large share of the tokens those jobs consume. And this week, Max and Team subscribers start receiving a monthly API credit for the Claude Platform, usable on any model: $100 a month on Max 5x, $200 on Max 20x and up to $500 on Team, pooled across the team. Anthropic is also adding beta computer-use and browser-use support to its Python and TypeScript SDKs, and names Haiku 5.5 as the model it suits best.
With Haiku 5.5, all three models Anthropic announced for the Claude 5.5 family are out within 15 days, and the bottom of the range has moved the most. Haiku 4.5 cost ten times what GPT-6 Luna does; Haiku 5.5 matches it. The benchmark lead over Luna is Anthropic's claim for now. The number to wait for is Artificial Analysis's cost per task, which will show whether the new tokenizer and the thinking Haiku 5.5 does by default eat into a sticker price that, per token, is already identical.
More on Claude
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.
