Anthropic released Claude Haiku 5.5 on 7 October 2026, and the headline price is the lowest in the Claude lineup: $0.10 per million input tokens and $0.50 per million output tokens. The footnote is the part worth reading. That rate holds for prompts up to 100,000 tokens. Past that, the prompt is priced at $0.50 and $2.50, and Haiku 5.5 is the only Claude model with a step like that.
The short answer
Haiku 5.5 (claude-haiku-5-5) bills $0.10 in and $0.50 out up to 100K prompt tokens, and five times that above it. Against Haiku 4.5's $1 and $5, the saving is closer to 7.7x than 10x for the same text, because the newer tokenizer counts about 30 percent more tokens. It also rejects non-default temperature, top_p and top_k with a 400 error.
What Anthropic published
The model ID is claude-haiku-5-5 on the Claude API, Google Cloud and Microsoft Foundry, and anthropic.claude-haiku-5-5 on Amazon Bedrock. It's also on Claude Platform on AWS. The documentation lists a 1M token context window, up to 128K output tokens, adaptive thinking with a default effort of medium, and a June 2026 knowledge cutoff. Anthropic calls it the fastest model in the current lineup. Retirement is set for no sooner than 7 October 2027, which gives you a year of floor.
Prices come in two bands. For prompts up to 100,000 tokens: $0.10 input, $0.50 output, $0.125 for a five-minute cache write, $0.20 for a one-hour write, and $0.01 for a cache read. Over 100,000 tokens: $0.50, $2.50, $0.625, $1 and $0.05. The Batch API takes 50 percent off both bands. Anthropic also cut Sonnet 5.5's cache-read price from $0.20 to $0.10 per million tokens the same day, according to its release notes.
The cliff, and what the tokenizer does to the saving
Every other Claude model since 4.6 includes the full 1M window at one standard rate. Anthropic's pricing page says so, and it names Haiku 5.5 as the exception: priced by prompt length. The arithmetic is blunt. A 100,000-token prompt costs $0.01 of input. A 100,001-token prompt costs $0.05, if the whole prompt moves to the higher rate. The page says a prompt over 100,000 tokens "pays higher prices" and doesn't spell out whether only the excess is repriced, so I'd budget the whole prompt at the higher rate until a bill says otherwise.
That doesn't make Haiku expensive. At $0.50 per million, a 900,000-token prompt costs $0.45 of input on Haiku 5.5 and $1.80 on Sonnet 5.5 at $2. The cliff matters when you compare Haiku with itself: the 1M window is real, but crossing 100K costs five times as much per token. If your retrieval or agent loop can stay under that line, it should.
The headline saving needs a correction too. Anthropic says Haiku 5.5 uses the tokenizer introduced with Claude 4.7, which turns the same text into about 30 percent more tokens than Haiku 4.5 did. So 1M tokens of old text is roughly 1.3M tokens now. At $0.10 that's $0.13, against $1 on Haiku 4.5, which is about 7.7 times cheaper. The 30 percent is an average and Anthropic notes it depends on content, so I might be wrong for your workload, especially with code or non-English text. Count a sample of your own prompts before you quote a number to anyone.
What to change before you switch
Three things in the documentation can break an existing integration. A non-default value for temperature, top_p or top_k returns a 400 error, so remove them rather than setting them to your old values. Thinking is adaptive and steered by the effort parameter, whose default here is medium. The manual budget_tokens mode is described as belonging to earlier models, so don't assume it carries over. And thinking blocks only work in the account that produced them, or one linked to it, which matters if you pass them between organisations.
Then re-test your prompts, because a cheaper model isn't the same model. Keep an eye on the cache: a read at $0.01 per million is where long, repeated prefixes get cheap, and it's worth restructuring prompts so the stable part comes first. If you track prices across vendors, the row is now in our LLM API pricing tracker, and the retirement floor is in the retirements calendar. Our earlier piece on Sonnet 5.5's price and thinking change covers the model one tier up.
Frequently asked questions
How much does Claude Haiku 5.5 cost?
$0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Above that the prices are $0.50 and $2.50. The Batch API halves both, and a cache read costs $0.01 or $0.05 per million depending on the band.
Does the whole prompt pay the higher rate above 100K tokens?
Anthropic prices the model by prompt length and says prompts over 100,000 tokens pay higher prices. It doesn't say whether only the excess is repriced. The safe reading is that the whole prompt moves to the higher rate, so plan for that.
Is Haiku 5.5 really 10 times cheaper than Haiku 4.5?
On the price list, yes: $0.10 against $1 per million input tokens. But the newer tokenizer produces about 30 percent more tokens for the same text, so the saving on identical content is nearer 7.7 times. Your own text may differ.
Where can I use Claude Haiku 5.5?
On the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. The ID is claude-haiku-5-5, with the anthropic. prefix on Bedrock.
Sources: Anthropic, Claude Haiku 5.5 announcement; the Claude Haiku 5.5 overview, pricing page and release notes, all read on 8 October 2026. The 7.7x figure and the 900,000-token comparison are our arithmetic from those prices.

