Live
Abstract illustration of a row of small violet token cubes running along a glowing line that steps sharply upward at one point, with a pink glow behind the step, on a deep purple gradient
AI & ML

Claude Haiku 5.5 Price: 90% Cheaper Until Your Prompt Hits 100K Tokens

Anthropic’s new small model costs a tenth of what its predecessor did, on paper. Claude Haiku 5.5, released Oct. 7, lists at US$0.10 per million input tokens (about CA$0.14) and US$0.50 per million output tokens (about CA$0.70) for prompts under 100,000 tokens. Haiku 4.5 charged US$1 and US$5.

Those base rates are identical to OpenAI’s GPT-6 Luna, the model OpenAI put into free ChatGPT. Two details in the fine print decide how much of the 90% cut a developer actually sees: a new tokenizer, and a price tier that jumps fivefold at 100,000 tokens.

Claude Haiku 5.5 pricing vs. Haiku 4.5 and GPT-6 Luna

Per million tokens Input Output Cache read
Haiku 5.5, prompts under 100K US$0.10 US$0.50 US$0.01
Haiku 5.5, prompts over 100K US$0.50 US$2.50 US$0.05
Haiku 4.5 US$1.00 US$5.00 US$0.10
GPT-6 Luna (base rate) US$0.10 US$0.50 n/a
Sonnet 5.5 US$2.00 US$10.00 US$0.10 (was US$0.20)

Prices are from VentureBeat’s launch report, which compiled Anthropic’s and competitors’ list prices. Batch jobs get a further 50% off, MarkTechPost reported. Haiku 5.5 has a 1M-token context window, up to 128,000 tokens of output, a June 2026 knowledge cutoff, and is available on the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry.

Anthropic paired the launch with two other price moves. Sonnet 5.5 cache reads fell by half to US$0.10, which Anthropic estimates saves about 20% on typical agent workloads. And Max and Team subscribers now get monthly API credits: US$100 (about CA$139) on Max 5x, US$200 on Max 20x and up to US$500 (about CA$695) shared across a Team plan, VentureBeat reported.

The tokenizer takes back part of the cut

Haiku 5.5 counts text differently. The new tokenizer “counts the same text as roughly 30% more tokens than Haiku 4.5,” according to MarkTechPost. You pay per token, so the same prompt now generates a bigger count before the discount applies.

Anthropic doesn’t hide this. Its own estimate, per VentureBeat, is that average workloads cost about 75% less than on Haiku 4.5 once the tokenizer is accounted for, against the 90% headline. Anthropic also says roughly 90% of Haiku 4.5 requests fall in the cheaper short-prompt tier. The other 10% are where the arithmetic turns.

The 100K-token cliff

Above 100,000 tokens, Haiku 5.5 costs five times as much per token. Now combine that with the tokenizer. A prompt that measured 80,000 tokens on Haiku 4.5 counts as roughly 104,000 on Haiku 5.5, if the 30% figure holds for your text, and lands in the expensive tier. Nothing about the prompt changed.

OpenAI’s tier break is further out. GPT-6 Luna’s higher rates start above 272,000 input tokens, so for a 150,000-token prompt Luna is cheaper on list price, MarkTechPost noted. For long-document work, retrieval over big knowledge bases or agents that drag long histories around, the price war is a lot less one-sided than the headline makes it look.

A Canadian startup’s monthly bill, worked out

Here is a hypothetical Toronto startup running two workloads. Token counts are measured on Haiku 4.5’s tokenizer, and we apply MarkTechPost’s 30% figure to both input and output for Haiku 5.5. Your own ratio will vary with your text.

Bar chart of monthly cost for two hypothetical workloads. Support bot with 4,000-token prompts: Haiku 4.5 US$6,500, Haiku 5.5 at the headline rate US$650, Haiku 5.5 with 30% more tokens US$845. Contract review with 85,000-token prompts: Haiku 4.5 US$4,500, Haiku 5.5 headline US$450, Haiku 5.5 with 30% more tokens US$2,925, because prompts cross into the 100K-token price tier.
The tokenizer barely dents the short-prompt saving but pushes 85,000-token prompts into Haiku 5.5’s pricier tier. Graphic: prompt/power

Workload 1: a support bot. One million requests a month, each with 4,000 input tokens and 500 output tokens.

  • Haiku 4.5: US$6,500 a month (about CA$9,035).
  • Haiku 5.5 at the headline rate, no tokenizer change: US$650.
  • Haiku 5.5 with 30% more tokens: US$845 (about CA$1,175), still 87% less than Haiku 4.5.

Workload 2: contract review. 50,000 requests a month, each sending an 85,000-token document and getting 1,000 tokens back.

  • Haiku 4.5: US$4,500 a month (about CA$6,255).
  • Haiku 5.5 if the prompt stayed under 100K: US$450.
  • Haiku 5.5 with 30% more tokens, which pushes each prompt to about 110,500 tokens and into the higher tier: US$2,925 (about CA$4,065), only 35% less than Haiku 4.5.

The support bot gets most of the promised saving. The contract reviewer gets about a third of it, and a bill six and a half times higher than the one the headline price implies.

What the benchmarks do and don’t say

On Anthropic’s own numbers, Haiku 5.5 scores 72.4% on the offline subset of OSWorld 2.1, a computer-use test, against 48.9% for GPT-6 Luna and 15.7% for Haiku 4.5. On Terminal-Bench 4.0 it scores 39.2% to Luna’s 16.4%. These are vendor-reported results, not independent tests.

The Terminal-Bench number comes with a condition. It was run at maximum effort, VentureBeat reported, while the default medium setting scores about 20%. Our read: higher effort generally means more reasoning tokens, and those are billed as output. The benchmark that makes Haiku look strongest is the configuration that costs the most to run.

There is a breaking change too. On Haiku 5.5, non-default temperature, top_p or top_k values return a 400 error, MarkTechPost reported. Code that sets temperature to 0 for reproducible output will fail until it is changed.

What to do before you switch

  • Re-count your real prompts with Haiku 5.5’s tokenizer before you budget. The 30% figure is an average from one outlet, not a constant.
  • Plot your prompt sizes. Anything between roughly 77,000 and 100,000 Haiku 4.5 tokens is at risk of crossing the line.
  • Strip sampling parameters from your API calls, or they will error.
  • Use caching and batch. Cache reads are US$0.01 per million in the short tier, and batch is half price.
  • Price long prompts against Luna as well as Haiku. Between 100,000 and 272,000 tokens, per-token list prices favour OpenAI.

Toronto’s Cohere sells token spending caps for AI agents, as we reported on Cohere’s North 2. A cap like that is only as good as the token count behind it. At 99,999 tokens, Haiku 5.5 bills US$0.10 per million. At 100,001, it is US$0.50.

// AI Editor
Cassandra Lee

Cassandra Lee covers AI and machine learning for prompt/power: the labs, the model releases, the research and the safety fights that come with them. She reads model cards the way other people read horoscopes: skeptically, and mostly for what's left unsaid.

Latest from prompt/power

  1. National Medal of Science Goes to Musk, Brin, Huang and Su. Nvidia Pledged US$1BOct 10
  2. TikTok Placebo Safety Test: New York Says Teens Got a Fake Feed ResetOct 10
  3. Waymo Takes Its First Loan, US$5 Billion, as Robotaxis Head OverseasOct 10
  4. Bell, Virgin, Public Mobile Plan Changes: The New Canadian Price ListOct 10
  5. Manus Raises More Than US$500 Million After China Blocked Meta’s DealOct 10

Leave a Reply

Your email address will not be published. Required fields are marked *