• About
  • Editorial policy
  • Contact
  • Privacy
  • Legal
  • Cookie settings
Sunday, October 11, 2026
  • Login
PacketNebula
  • Home
  • News
  • Tools
    • Network
    • Security
    • Developer
    • Sysadmin
    • SEO
    • Email & DNS
  • Guides
  • Trackers
    • LLM API pricing
    • Open-weight models
    • Retirements calendar
    • Infrastructure deals
  • About
  • Download
No Result
View All Result
PacketNebula
No Result
View All Result
Home Dev

Claude Haiku 5.5 is $0.10 per million until a prompt passes 100K

by Stéphane Cardon
8 October 2026
in Dev
0
Anthropic's announcement graphic for Claude Haiku 5.5: the model name set in serif type across torn paper collage panels

Image: Anthropic

Share on FacebookShare on Twitter

Anthropic released Claude Haiku 5.5 on 7 October 2026, and the headline price is the lowest in the Claude lineup: $0.10 per million input tokens and $0.50 per million output tokens. The footnote is the part worth reading. That rate holds for prompts up to 100,000 tokens. Past that, the prompt is priced at $0.50 and $2.50, and Haiku 5.5 is the only Claude model with a step like that.

The short answer

Haiku 5.5 (claude-haiku-5-5) bills $0.10 in and $0.50 out up to 100K prompt tokens, and five times that above it. Against Haiku 4.5's $1 and $5, the saving is closer to 7.7x than 10x for the same text, because the newer tokenizer counts about 30 percent more tokens. It also rejects non-default temperature, top_p and top_k with a 400 error.

$0.10 / $0.50per million tokens, prompts up to 100K
5xthe rate above 100K prompt tokens
about 7.7xcheaper than Haiku 4.5, same text

What Anthropic published

The model ID is claude-haiku-5-5 on the Claude API, Google Cloud and Microsoft Foundry, and anthropic.claude-haiku-5-5 on Amazon Bedrock. It's also on Claude Platform on AWS. The documentation lists a 1M token context window, up to 128K output tokens, adaptive thinking with a default effort of medium, and a June 2026 knowledge cutoff. Anthropic calls it the fastest model in the current lineup. Retirement is set for no sooner than 7 October 2027, which gives you a year of floor.

Prices come in two bands. For prompts up to 100,000 tokens: $0.10 input, $0.50 output, $0.125 for a five-minute cache write, $0.20 for a one-hour write, and $0.01 for a cache read. Over 100,000 tokens: $0.50, $2.50, $0.625, $1 and $0.05. The Batch API takes 50 percent off both bands. Anthropic also cut Sonnet 5.5's cache-read price from $0.20 to $0.10 per million tokens the same day, according to its release notes.

The cliff, and what the tokenizer does to the saving

Every other Claude model since 4.6 includes the full 1M window at one standard rate. Anthropic's pricing page says so, and it names Haiku 5.5 as the exception: priced by prompt length. The arithmetic is blunt. A 100,000-token prompt costs $0.01 of input. A 100,001-token prompt costs $0.05, if the whole prompt moves to the higher rate. The page says a prompt over 100,000 tokens "pays higher prices" and doesn't spell out whether only the excess is repriced, so I'd budget the whole prompt at the higher rate until a bill says otherwise.

That doesn't make Haiku expensive. At $0.50 per million, a 900,000-token prompt costs $0.45 of input on Haiku 5.5 and $1.80 on Sonnet 5.5 at $2. The cliff matters when you compare Haiku with itself: the 1M window is real, but crossing 100K costs five times as much per token. If your retrieval or agent loop can stay under that line, it should.

The headline saving needs a correction too. Anthropic says Haiku 5.5 uses the tokenizer introduced with Claude 4.7, which turns the same text into about 30 percent more tokens than Haiku 4.5 did. So 1M tokens of old text is roughly 1.3M tokens now. At $0.10 that's $0.13, against $1 on Haiku 4.5, which is about 7.7 times cheaper. The 30 percent is an average and Anthropic notes it depends on content, so I might be wrong for your workload, especially with code or non-English text. Count a sample of your own prompts before you quote a number to anyone.

What to change before you switch

Three things in the documentation can break an existing integration. A non-default value for temperature, top_p or top_k returns a 400 error, so remove them rather than setting them to your old values. Thinking is adaptive and steered by the effort parameter, whose default here is medium. The manual budget_tokens mode is described as belonging to earlier models, so don't assume it carries over. And thinking blocks only work in the account that produced them, or one linked to it, which matters if you pass them between organisations.

Then re-test your prompts, because a cheaper model isn't the same model. Keep an eye on the cache: a read at $0.01 per million is where long, repeated prefixes get cheap, and it's worth restructuring prompts so the stable part comes first. If you track prices across vendors, the row is now in our LLM API pricing tracker, and the retirement floor is in the retirements calendar. Our earlier piece on Sonnet 5.5's price and thinking change covers the model one tier up.

Frequently asked questions

How much does Claude Haiku 5.5 cost?

$0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Above that the prices are $0.50 and $2.50. The Batch API halves both, and a cache read costs $0.01 or $0.05 per million depending on the band.

Does the whole prompt pay the higher rate above 100K tokens?

Anthropic prices the model by prompt length and says prompts over 100,000 tokens pay higher prices. It doesn't say whether only the excess is repriced. The safe reading is that the whole prompt moves to the higher rate, so plan for that.

Is Haiku 5.5 really 10 times cheaper than Haiku 4.5?

On the price list, yes: $0.10 against $1 per million input tokens. But the newer tokenizer produces about 30 percent more tokens for the same text, so the saving on identical content is nearer 7.7 times. Your own text may differ.

Where can I use Claude Haiku 5.5?

On the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. The ID is claude-haiku-5-5, with the anthropic. prefix on Bedrock.

Sources: Anthropic, Claude Haiku 5.5 announcement; the Claude Haiku 5.5 overview, pricing page and release notes, all read on 8 October 2026. The 7.7x figure and the 900,000-token comparison are our arithmetic from those prices.

Tags: anthropicAPI pricingclaudenews
ShareTweet
Previous Post

Reflection Beam is a 501B open model with no weights yet

Next Post

GPT-6.1 Sol Ultrafast bills six times Sol and 1.2 times Astra

Stéphane Cardon

Stéphane Cardon

Network cybersecurity engineer: architecture, LAN and WLAN, system administration. He runs PacketNebula, where he publishes the tools and answers he wanted for his own work.

Next Post
OpenAI's announcement video still for Ultrafast: a presenter at a desk with a laptop and the caption 'Ultrafast, now in the API, ChatGPT Work, and Codex'

GPT-6.1 Sol Ultrafast bills six times Sol and 1.2 times Astra

Free tools

  • IPv4 subnet calculator
  • DNS lookup
  • SSL certificate checker
  • HTTP headers checker
  • SPF, DKIM and DMARC checker
  • JWT decoder
  • Cron expression builder
  • Chmod calculator
  • Password generator

All 24 tools →

Guides to keep handy

  • Flush the DNS cache on Windows
  • Kill the process using a port
  • Generate an SSH key
  • Secure a new Ubuntu VPS
  • Find files on Linux
  • Extract and create tar.gz

More guides →

Who writes this

Stéphane Cardon, network cybersecurity engineer. Architecture, LAN and WLAN, system administration. PacketNebula holds the tools and answers he wanted for his own work.

About this site →

  • About
  • Editorial policy
  • Contact
  • Privacy
  • Legal
  • Cookie settings

Copyright © 2026 Stephane Cardon.

No Result
View All Result
  • Home
  • News
  • Tools
    • Network
    • Security
    • Developer
    • Sysadmin
    • SEO
    • Email & DNS
  • Guides
  • Trackers
    • LLM API pricing
    • Open-weight models
    • Retirements calendar
    • Infrastructure deals
  • About
  • Download

Copyright © 2026 Stephane Cardon.

Cookies, your call

We use Google Analytics to see which pages get read. Once our AdSense application is approved, articles and guides will carry Google ads. Nothing runs until you say yes, and nothing runs on the tool pages. Details