Models·8 min read
By BitsMindsSource: Anthropic

Claude Haiku 5.5 Is Out at a Tenth of Haiku 4.5's Price

Anthropic released Claude Haiku 5.5 on 7 October at $0.10 and $0.50 per million tokens for prompts up to 100,000 tokens, a tenth of Haiku 4.5's price. That is exactly what GPT-6 Luna costs, and on Anthropic's own benchmarks Haiku 5.5 beats Luna on every test the two share.

Claude Haiku 5.5: Anthropic's fastest model, at a tenth of Haiku 4.5's price A terracotta and ivory stopwatch, tipped to show its depth, has the official Claude asterisk inlaid in clay as the hub of its green sweep hand. A paper tag tied to the crown reads minus 90 percent against Haiku 4.5. Beside it, the title names Claude Haiku 5.5 and quotes its API prices for prompts up to 100,000 tokens: 10 US cents per million input tokens and 50 cents per million output tokens. The stopwatch is an editorial metaphor for Anthropic's claim that Haiku 5.5 is its fastest model at standard speed, not an Anthropic product; the 90 percent figure applies to prompts up to 100,000 tokens, and longer prompts cost five times as much. 51015202530354045505560 CLAUDE HAIKU 5.5 PER-TOKEN PRICE −90% VS HAIKU 4.5 PROMPTS UP TO 100K CLAUDE Haiku 5.5 $0.10 INPUT $0.50 OUTPUT PER MILLION TOKENS BITSMINDS.COM
Share:

Claude Haiku 5.5 is Anthropic's new small model, released on 7 October 2026 the third Claude 5.5 model, after Opus 5.5 on 22 September and Sonnet 5.5 on 28 September. Its model ID is claude-haiku-5-5. It has a 1M-token context window and up to 128K tokens of output, up from 200K and 64K on Haiku 4.5, and costs $0.10 per million input tokens and $0.50 per million output for prompts up to 100,000 tokens. It is available now on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS.

Anthropic calls it the cheapest, fastest and most capable small model it has released. The jobs it names are high-volume and repetitive: summaries, context compaction, database queries and classification, plus work as a subagent under Opus 5.5 or Sonnet 5.5 on coding tasks. Speed is the other half of the pitch. Haiku 5.5 is Anthropic's fastest model at standard speed, though a footnote concedes that Opus models in Fast Mode are quicker, and Anthropic points to live customer support and browser automation as the work where that speed pays.

GPT-6 Luna's price, to the cent

The price lands exactly on OpenAI's. GPT-6 Luna, OpenAI's budget model since 22 September, also costs $0.10 per million input tokens and $0.50 per million output, with cached input at $0.01, and Haiku 5.5's cache reads cost the same $0.01 for prompts under the 100K mark. Anthropic put Luna in its own comparison, and on its own table Haiku 5.5 is ahead on every test where both have a score.

The next three charts are Anthropic's figures, run and selected by the company that sells the model. No independent numbers exist yet: Artificial Analysis had not published Haiku 5.5 when this was written. On its Intelligence Index, version 4.3.2, GPT-6 Luna at maximum effort scores 38 and costs $0.07 per index task, the bar Haiku 5.5 will have to clear on cost per task rather than on price per token.

Computer use and agentic codingAccuracy, % — higher is better · Anthropic's own figures, 7 October 2026Claude Haiku 5.5Claude Haiku 4.5GPT-6 LunaClaude Sonnet 5.502040608010072.415.748.983.9OSWorld 2.1(offline subset)39.20.016.470.6Terminal-Bench4.046.4n/r42.452.1FrontierCode 1.1
Haiku 4.5 scores zero on Terminal-Bench 4.0 and has no FrontierCode result. Sonnet 5.5's FrontierCode bar is at Xhigh effort; its OSWorld figure here is 83.9%, where its own launch post gave 80.1%. Source: Anthropic.

Computer use is the widest gap. On the offline subset of OSWorld 2.1, where an agent operates a real computer through long multi-step tasks, Haiku 5.5 scores 72.4% to Luna's 48.9% and Haiku 4.5's 15.7%, within about 12 points of Sonnet 5.5. Terminal-Bench 4.0 shows the limit. Haiku 5.5 more than doubles Luna there at 39.2%, but Sonnet 5.5 reaches 70.6%, and Anthropic says plainly that Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding. FrontierCode, which asks whether a change could be merged without human edits, is closer: 46.4% against Luna's 42.4%.

Reasoning and chart readingAccuracy, % — higher is better · Chartography without tools · Anthropic's own figuresClaude Haiku 5.5Claude Haiku 4.5GPT-6 LunaClaude Sonnet 5.502040608010045.910.2n/r56.9HLE, no tools57.418.7n/r64.5HLE, with tools46.46.429.161.6Chartography
Anthropic gives no Humanity's Last Exam score for GPT-6 Luna. Against Haiku 4.5, Haiku 5.5 scores between three and seven times higher on all three tests.

The reasoning and vision rows show how far a single generation moved the small model. Haiku 4.5 scored 6.4% on Chartography; Haiku 5.5 scores 46.4%, against Luna's 29.1%. On Humanity's Last Exam, a test of expert-level academic knowledge, it gets 45.9% without tools and 57.4% with them.

Knowledge work, EloAxis from zero · run by Artificial Analysis, published by AnthropicClaude Haiku 5.5Claude Haiku 4.5GPT-6 LunaClaude Sonnet 5.50500100015002000162073514371840GDPval-AA v2.1157861413361824AA-Briefcasev1.1
On both professional-task boards, Haiku 5.5 lands 180 to 250 points ahead of GPT-6 Luna, a similar distance behind Sonnet 5.5, and roughly 900 points above Haiku 4.5.

The Elo boards come from Artificial Analysis and rate professional work; GDPval-AA, for instance, draws its tasks from 44 occupations. Haiku 5.5 scores 1620 on it, 183 points ahead of Luna and 220 behind Sonnet 5.5.

Where the 75% comes from

Anthropic's headline is that Haiku 5.5 costs about 75% less to run than Haiku 4.5. The per-token cut is larger and the per-task figure is smaller, for two reasons spelled out in the footnotes. First, the price depends on prompt length. Up to 100,000 tokens, every rate is a tenth of Haiku 4.5's. Above that, input costs $0.50 and output $2.50, still half of Haiku 4.5's $1 and $5, but five times Haiku 5.5's own short-prompt rate. Anthropic says 90% of requests to Haiku 4.5 fell under the line. Second, Haiku 5.5 uses the newer tokenizer from Sonnet 5.5 and Opus 5.5. The announcement says this means slightly more tokens per task; the developer docs put it at about 30% more tokens for the same text than Haiku 4.5.

Per million tokensInputCache readOutput
Claude Haiku 5.5, prompts up to 100K$0.10$0.01$0.50
Claude Haiku 5.5, prompts over 100K$0.50$0.05$2.50
GPT-6 Luna$0.10$0.01$0.50
Claude Haiku 4.5$1.00$0.10$5.00
Claude Sonnet 5.5$2.00$0.10$10.00

The tiered price is new in the current lineup. Anthropic's pricing page says every model from Claude 4.6 on bills a 900K-token request at the same rate as a 9K one, with Haiku 5.5 named as the exception. Cache writes follow the same split: $0.125 for five minutes and $0.20 for an hour under the line, $0.625 and $1 above it. The Batch API halves both tiers and accepts up to 300K output tokens with a beta header. For classification, routing and short summaries, the 100K line will rarely matter. For retrieval jobs that stuff long documents into each call, it decides the bill.

Effort settings reach Haiku

Haiku 5.5 is the first Haiku with an adjustable effort setting, from Low to Max, so developers can trade answer quality against speed and cost as they can on the larger models. It defaults to Medium. Adaptive thinking is on by default, but unlike Opus 5.5, Haiku 5.5 still lets a request turn thinking off, at High effort or below. Anthropic's launch post plots each effort setting against cost per attempt on OSWorld, GDPval-AA and Humanity's Last Exam, but it does not print the values behind the dots.

Moving from Haiku 4.5 is not a model-name swap. Anthropic's list of changes includes five that break existing code:

  • Manual extended thinking returns an error. budget_tokens is gone; adaptive thinking and the effort parameter replace it.
  • Non-default sampling parameters return an error. Requests have to omit temperature, top_p and top_k.
  • Prefilling the assistant's reply returns an error. Every request has to end with a user turn.
  • Computer use needs the new toolset, computer_toolset_20260801 in place of computer_20250124, on the Claude API and Google Cloud.
  • Editing an earlier turn invalidates thinking blocks. Conversations that send thinking blocks back have to stay append-only.

Quieter changes can still trip a working app. A response can now begin with a thinking block, so code that reads the first content block as the answer has to select blocks by type. Thinking text is omitted unless a request asks for a summary. Safety classifiers can decline a request with a refusal stop reason, with no server-side fallback. Thinking tokens count toward max_tokens, so a small limit can end a response before any text appears. The migration guide covers each change.

What early customers reported

Anthropic published results from six companies that tested the model before launch, each on its own workload. HubSpot says Haiku 5.5 posted the best score it has seen on its CRM evaluation suite, 92.8% averaged over three runs. AlphaSense ran 400 queries through Ask in Document, a feature it says handles about 8 million calls a week, and measured 0.84 against Haiku 4.5's 0.76. Box reports a score 11 points above Haiku 4.5 at about half the latency, and Asana a cut of more than 30% in task-completion latency against the model it uses today. Cognition says its Devin Fusion setup, with Opus 5.5 leading and Haiku 5.5 as the helper, holds a FrontierCode score of 66.2 at lower cost and latency. These are the customers Anthropic chose to quote, on tests they designed.

Safeguards: between Haiku 4.5 and Sonnet 5.5

Anthropic reports large improvements over Haiku 4.5 on almost all of its alignment evaluations, with far fewer cases of misaligned behaviour and less willingness to help with misuse; the system card has the detail. Its cybersecurity safeguards are stricter than Haiku 4.5's but looser than those on Anthropic's other recent models: they allow a wider range of defensive work than Sonnet 5.5's, while still blocking penetration testing and other techniques more often used by attackers. Biology safeguards match Sonnet 5, Sonnet 5.5 and Opus 5.

Also announced: cheaper Sonnet 5.5 caching and API credit for subscribers

Anthropic paired the launch with two price moves. Cache reads on Sonnet 5.5 are halved from $0.20 to $0.10 per million tokens from 7 October, which Anthropic estimates makes Sonnet 5.5 about 20% cheaper on most agentic work, because cached context makes up a large share of the tokens those jobs consume. And this week, Max and Team subscribers start receiving a monthly API credit for the Claude Platform, usable on any model: $100 a month on Max 5x, $200 on Max 20x and up to $500 on Team, pooled across the team. Anthropic is also adding beta computer-use and browser-use support to its Python and TypeScript SDKs, and names Haiku 5.5 as the model it suits best.

With Haiku 5.5, all three models Anthropic announced for the Claude 5.5 family are out within 15 days, and the bottom of the range has moved the most. Haiku 4.5 cost ten times what GPT-6 Luna does; Haiku 5.5 matches it. The benchmark lead over Luna is Anthropic's claim for now. The number to wait for is Artificial Analysis's cost per task, which will show whether the new tokenizer and the thinking Haiku 5.5 does by default eat into a sticker price that, per token, is already identical.

More on Claude

Evergreen coverage we keep current — start here.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Reflection Beam: a large core, a narrow active path An original monochrome optical sculpture on a dark reflective laboratory bench. A narrow white beam passes through a hollow silver ring and branches toward a few illuminated cartridges within a large graphite core. Most cartridges remain dark; their paths recombine into a beam that reaches a solid circular receiver. The hollow and solid circles echo Reflection's mark. The labels read BEAM, 501B total parameters, 23B active per token, PREVIEW and WEIGHTS PLANNED — OCT. The illustration represents sparse inference in the model, not physical hardware or its exact expert layout. BitsMinds original editorial vector artwork for reflection-ai-beam-501b-open-weight-model. Facts verified on 6 October 2026 at https://reflection.ai/blog/introducing-beam: 501B total parameters, 23B active, weights planned later in October under Apache 2.0. The module layout is symbolic. No performance ranking or measured-cost claim is depicted. Self-contained SVG with the project Reflection wordmark. BEAM 501B TOTAL PARAMETERS 23B ACTIVE PER TOKEN PREVIEW WEIGHTS PLANNED · OCT BITSMINDS.COM
Models

Reflection’s Beam: A 501B Open Model Built to Be Frugal

Opus 5.5 vs GPT-6 Astra vs Grok 4.7 Build No Man's Sky
Models

Opus 5.5 vs GPT-6 Astra vs Grok 4.7 Build No Man's Sky

Claude Fable 5.5: the leaked videos A dark film slate with the Claude symbol inlaid in terracotta stands in front of a strip of film frames on a warm cream background. The slate reads Fable 5.5, and its fields read Roll: routed, Scene: Tuesday?, Take: unconfirmed. The image illustrates an unconfirmed rumour, not an announced model. FABLE 5.5 ROLL ROUTED SCENE TUESDAY? TAKE UNCONFIRMED BITSMINDS.COM
Models

Claude Fable 5.5 Leaks: Viral Videos and a Tuesday Rumour