• Latest
  • Trending
  • All
Answer card for Claude Opus 5.5: 4 dollars in and 20 dollars out per million tokens, down from 5 and 25, with cache reads at 20 cents.

Claude Opus 5.5 drops to $4 and $20, and breaks four Opus 5 habits

23 September 2026
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Wednesday, September 23, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Claude Opus 5.5 drops to $4 and $20, and breaks four Opus 5 habits

by stephane
23 September 2026
in Dev
0
Answer card for Claude Opus 5.5: 4 dollars in and 20 dollars out per million tokens, down from 5 and 25, with cache reads at 20 cents.
491
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Swap the model string on an agent that forces a tool call, and you'll get a 400 before the first token. That's the Claude Opus 5.5 launch in one line. The new model, claude-opus-5-5, went live on 22 September at $4 per million input tokens and $20 per million output, down from Opus 5's $5 and $25, and Anthropic's own table puts it ahead of Fable 5.1 on most rows. It also refuses four request shapes that Opus 5 accepted, and the cheapest bill in the world doesn't help a pipeline that won't start.

The short answer

Claude Opus 5.5 shipped 22 September 2026 at $4 and $20 per million tokens, cache reads at $0.20, with the same 1M context and 128K output as Opus 5. Anthropic's table has it beating Fable 5.1 on Terminal-Bench 4.0 by more than ten points. Before you switch, remove any disabled thinking or forced tool_choice, and set effort explicitly, because the default fell from high to medium.

$4 / $20per 1M, down from $5 / $25
66.4%Terminal-Bench 4.0, Fable 5.1 at 55.8%
4breaking changes for Opus 5 code
Answer card for Claude Opus 5.5: the model id claude-opus-5-5 costs 4 dollars in and 20 dollars out per million tokens, down from 5 and 25, with cache reads at 20 cents and the same 1M window.
Pricing and limits as listed on Anthropic's model page on 23 September 2026.

The price cut is bigger on cache than on output

Twenty percent off the headline. That's the part everyone quoted. The line we'd actually watch is cache reads: $0.20 per million, which is 5% of the base input price. On Opus 5 a cache read was 10% of $5, so fifty cents. A 60% cut. If you run long agent sessions that re-read the same 200K of repository context every turn, that's where the invoice moves, not on output.

The rest of the card is plainer. Five minute cache writes cost $5 per million, one hour writes $8, and batch halves everything to $2 and $10. The context window is still 1M tokens, max output is still 128K on the synchronous API, and batch can go to 300K with the output-300k-2026-03-24 beta header. Knowledge cutoff is June 2026. Retirement is "not sooner than" 22 September 2027, so you've got a year of guaranteed runway if you pin it.

Anthropic also claims the model spends about 40% less than Opus 5 on typical workloads, partly because it writes shorter answers. We can't check that yet. It's the vendor's number on the vendor's mix, and it runs straight into a second fact from the same docs: at a given effort level, Opus 5.5 thinks more per turn than Opus 5, "most of all at xhigh and max". Thinking tokens bill as output. So a cheaper rate and a hungrier model can net out anywhere, and I honestly don't know where they'll land on a real coding loop until someone publishes token counts per task.

One ugly data point already exists. Simon Willison ran his usual SVG test at max effort and the model burned through the whole 128K output budget mid reasoning, twice, at $2.56 a run and close to 20 minutes each. Fable 5.1 finished the same prompt. It's one prompt. It's also the kind of failure you don't want to discover on a nightly batch.

Four request shapes that now return a 400

Here's the list from Anthropic's migration notes, in the order we'd check a codebase. First, thinking can't be disabled. Opus 5 accepted thinking: {"type": "disabled"} at high effort or below. Opus 5.5 answers that, and any manual budget_tokens, with an invalid_request_error. Drop the field or send {"type": "adaptive"}, then use effort as the dial.

Second, forced tool use is gone. tool_choice set to any or to a named tool fails, including on the token counting endpoint. We wrote about the same break on Fable 5.1 three weeks ago, and it's the one that bites structured extraction jobs hardest. The documented fix is auto plus strict tool use, or moving the schema to structured outputs.

Third, the older computer_20251124 tool is refused on the Claude API and Google Cloud. You need the computer_toolset_20260801 toolset there. Bedrock still accepts the old tool, which is a nice trap if you test on one cloud and deploy on another. Fourth, thinking blocks are bound to the model and the conversation. Opus 5.5 reads blocks from Opus 5, but not from Fable or Mythos, and on accounts created from 31 August 2026 a replayed block after an edited system prompt or tool list returns a 400.

Then there's the change that doesn't error at all, which we think is the nastiest. The short notes the model writes between tool calls now arrive as thinking blocks, and at the default display: "omitted" their text is empty. If your UI streams those notes as progress updates, it just goes silent. No exception, no log line. A user thinks the agent froze.

Checklist of Opus 5 code on Opus 5.5: the 1M context, 128K output, batch at 2 and 10 dollars, adaptive thinking code and computer_20251124 on Bedrock carry over, while disabled thinking, forced tool_choice and computer_20251124 on the API return a 400, and default effort drops to medium.
Built from the "What's new in Claude Opus 5.5" page. The red column is what fails, or goes quiet, without a code change.

The benchmark table, and where we'd use it

Anthropic's launch page puts Opus 5.5 at 66.4% on Terminal-Bench 4.0, against 55.8% for Fable 5.1 and 52.3% for Opus 5. That's a mid tier model beating the top one by more than ten points on long terminal work, at 40% of Fable's $10 and $50 list price. Other rows are closer: 57.8% on CursorBench 4.0 against Fable's 51.8%, 54.4% on FrontierCode v1.1 against 50.3%. On AutomationBench it scores 40.0%, a hair behind the 41.4% the same page gives GPT-6 Astra.

Two cautions. The page doesn't say which effort level produced each row, and the default is now medium, so your out of the box result may not look like the table. And harnesses disagree: xAI's Grok 4.7 page listed Fable 5.1 at 57.9% on Terminal-Bench 4.0, two points above Anthropic's own figure. No independent run had been published when we wrote this.

Bar chart of Terminal-Bench 4.0 scores from Anthropic's Opus 5.5 page: Claude Opus 5.5 at 66.4 percent, Claude Fable 5.1 at 55.8 percent and Claude Opus 5 at 52.3 percent.
Vendor figures from the launch page, not yet reproduced independently.

Our take: if you're on Opus 5 with thinking already on and no forced tools, move now, pin effort to whatever you ran before, and compare a week of invoices. If you disabled thinking to save money, the math changes, because that option is gone and low effort is the closest substitute. If you're on Fable 5.1 for agent work, we'd run a side by side before paying 2.5x for it again. I'd keep Fable for the hardest one shot tasks for now, mostly because of that max effort result. Sonnet 5.5 and Haiku 5.5 are due "in the coming weeks" per TechCrunch, so the cheap end of the lineup is about to move too.

A quick check that your key can see the model before you touch a config:

Linux
curl -s https://api.anthropic.com/v1/models/claude-opus-5-5 -H "x-api-key: $ANTHROPIC_API_KEY" -H "anthropic-version: 2023-06-01"

Sources

Release date, pricing, cache rates, context, output limits, cutoff, retirement date and availability come from Anthropic's Claude Opus 5.5 model page. The four breaking changes, the silent progress update change, the default effort and the extra thinking per turn are from What's new in Claude Opus 5.5. Benchmark scores and the 40% workload claim are from the Opus 5.5 launch page. The previous $25 output price and the coming Sonnet and Haiku releases are reported by TechCrunch. The max effort test and the 60% cache cut are from Simon Willison. The figures are ours, built from those pages.

Frequently asked questions

How much does Claude Opus 5.5 cost?

$4 per million input tokens and $20 per million output tokens, with cache reads at $0.20, five minute cache writes at $5 and one hour writes at $8. Batch requests are half price, so $2 and $10. Those are Anthropic's list prices on 23 September 2026.

Can I turn off thinking on Opus 5.5?

No. Adaptive thinking is always on, and a request with thinking disabled or a manual token budget returns a 400 error. Lower the effort parameter instead. The lowest setting is the closest thing to the old behaviour, but it isn't identical.

Is Opus 5.5 better than Fable 5.1?

On Anthropic's own table it scores higher on most rows, including Terminal-Bench 4.0 at 66.4% against 55.8%, for 40% of the price. Those are vendor numbers, and one early independent test at max effort failed where Fable 5.1 succeeded. We'd test both on your own tasks before retiring Fable.

Why did my agent's progress messages stop after switching?

Opus 5.5 returns the text it writes between tool calls as thinking blocks, and at the default display setting that text is empty. Nothing errors. Set a thinking.display value that returns the text, as described in Anthropic's migration guide.

When will Claude Opus 5 be retired?

We didn't find a retirement date for Opus 5 in the launch material or the Opus 5.5 docs, so there's no deadline forcing the move yet. Opus 5.5 itself is guaranteed until at least 22 September 2027, per its model page.

Tags: AI pricinganthropicclaudellmnews
Share196Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card for Claude Opus 5.5: 4 dollars in and 20 dollars out per million tokens, down from 5 and 25, with cache reads at 20 cents.

Claude Opus 5.5 drops to $4 and $20, and breaks four Opus 5 habits

23 September 2026
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.