• Latest
  • Trending
  • All
Answer card for Claude Sonnet 5.5 showing 2 dollars in and 10 dollars out per million tokens, unchanged from Sonnet 5, and between_tools replacing disabled thinking.

Claude Sonnet 5.5 keeps $2 and $10 and retires thinking disabled

29 September 2026
Answer card: Reflection AI announced Beam, a 501B parameter open-weight mixture of experts model, with weights promised later in October 2026.

Reflection Beam is a 501B open model with no weights yet

6 October 2026
Answer card: Aleph Alpha released Kolibri-1 on 3 October 2026 as an Apache 2.0 mixture of experts model with 78.1 billion total and 3.46 billion active parameters, served at 1 million tokens of context but trained on sequences of up to 256 thousand tokens.

Kolibri-1 serves 1M tokens but was trained to 256K

4 October 2026
Answer card on Gemini 4 Argon pricing and gated access, announced by Google on 30 September 2026.

Gemini 4 Argon costs $2 and $10 now, $4 and $20 later

2 October 2026
Answer card for GPT-6.1 Sol showing 2 dollars in and 10 dollars out per million tokens, unchanged from GPT-6 Sol, with cached input at 10 cents.

GPT-6.1 Sol costs what GPT-6 Sol did, and only the cache got cheaper

30 September 2026
Answer card: Xiaomi retrained MiMo-V2.6 to stop repeating tool calls, API swapped on 25 September 2026, MIT weights on 27 September, same model names.

MiMo-V2.6-Pro was quietly retrained to stop looping on tool calls

28 September 2026
Answer card: Anthropic committed $11.6 billion over seven years to Akamai Cloud for CPU workloads only, with revenue from the second half of 2027.

Anthropic’s $11.6B Akamai deal buys CPUs, not GPUs

27 September 2026
Answer card: Google Suncatcher MVP satellite, four Trillium TPUs on about one kilowatt, launching on SpaceX Transporter-18, reported for 1 October 2026.

Google’s first Suncatcher satellite flies four TPUs on 1 kW

25 September 2026
Answer card for Claude Opus 5.5: 4 dollars in and 20 dollars out per million tokens, down from 5 and 25, with cache reads at 20 cents.

Claude Opus 5.5 drops to $4 and $20, and breaks four Opus 5 habits

23 September 2026
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Wednesday, October 7, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Claude Sonnet 5.5 keeps $2 and $10 and retires thinking disabled

by stephane
29 September 2026
in Dev
0
Answer card for Claude Sonnet 5.5 showing 2 dollars in and 10 dollars out per million tokens, unchanged from Sonnet 5, and between_tools replacing disabled thinking.
494
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

If your Sonnet 5 code sends thinking: {"type": "disabled"}, Claude Sonnet 5.5 will hand it back as a 400. That's the first thing we'd check. The model itself went live on 28 September as claude-sonnet-5-5, at the same $2 per million input tokens and $10 per million output as Sonnet 5, and on Anthropic's own table it sits within two or three points of Opus 5.5, a model that costs twice as much. Cheap and close to the top. Also a bit stricter about how you call it.

The short answer

Claude Sonnet 5.5 shipped 28 September 2026 at $2 and $10 per million tokens, cache reads at $0.20, with a 1M context window and 128K output. It scores 70.6% on Terminal-Bench 4.0, ahead of Opus 5.5, and trails it by a couple of points elsewhere. Before switching, replace disabled thinking with between_tools and drop forced tool_choice.

$2 / $10per 1M, same as Sonnet 5
70.6%Terminal-Bench 4.0, per Anthropic
5breaking changes for Sonnet 5 code
Answer card for Claude Sonnet 5.5: the model id claude-sonnet-5-5 costs 2 dollars in and 10 dollars out per million tokens, unchanged from Sonnet 5, with cache reads at 20 cents, a 1M window, and between_tools replacing disabled thinking.
Pricing and limits as listed on Anthropic's model page on 29 September 2026.

The cheap model is now close to the expensive one

Nothing moved on the price card. Input is $2, output $10, five minute cache writes $2.50, one hour writes $4, cache reads $0.20, and batch halves all of it. The window is 1M tokens, max output is 128K on the synchronous API, and batch can go to 300K with the output-300k-2026-03-24 beta header. Knowledge cutoff is June 2026. Retirement is "not sooner than" 28 September 2027.

What moved is the gap to Opus 5.5. On Anthropic's published table Sonnet 5.5 gets 55.5% on CursorBench 4.0 against 57.8%, 80.1% on OSWorld 2.1 against 81.8%, and 1844 Elo on GDPval-AA against 1846. Two Elo points. On Terminal-Bench 4.0 it actually leads, 70.6% against 66.4% for Opus 5.5 at xhigh effort. For a model at half the per token price, that's a strange place for the flagship to be.

Here's where I'd slow down. The same table gives Sonnet 5 just 10.3% on Terminal-Bench 4.0, which would make this a sevenfold jump in one point release. Maybe it is. I'd bet more on a harness or setting that Sonnet 5 handled badly, and the page doesn't explain it. These are vendor numbers, and Artificial Analysis flagged that it tested a pre-release build with a structured output bug. There's also a footnote we liked for its honesty: on FrontierCode, max effort scored lower than xhigh (46.2% against 52.1%), because more effort sometimes sent the model into code review loops that ended in timeouts or out of scope edits.

The efficiency claim is the one that matters for a bill. Anthropic says output is "30%+ faster" and a task costs "up to 30% less", mostly through fewer tokens and tool calls. GitHub's changelog says the same thing from its side: it matched Sonnet 5 on coding tasks with "significantly fewer steps". Up to. We'd still measure it on our own traces before believing the invoice.

Bar chart from Anthropic's table: on Terminal-Bench 4.0 Sonnet 5.5 scores 70.6 percent against 66.4 for Opus 5.5 at xhigh, and on CursorBench 4.0 Opus 5.5 scores 57.8 against 55.5 for Sonnet 5.5.
Vendor scores from the Sonnet 5.5 launch page, 28 September 2026. Our chart.

between_tools, and the other four breaks

If you read our Opus 5.5 piece last week, most of this will look familiar. Sonnet gets the same treatment with one twist. Thinking still has an off switch, it's just renamed and narrower. You send thinking: {"type": "between_tools"}, which kills the up-front reasoning but keeps the short progress notes between tool calls. It only works at low, medium or high effort. At xhigh or max it's a 400, it won't take display or budget_tokens, and effort can't change mid-conversation while it's on. Without tools, you get plain text back, like disabled did.

The other four. Forced tool use is gone: tool_choice of any or a named tool returns a 400, even on token counting, so structured extraction jobs need auto plus strict tools or structured outputs (the same break Fable 5.1 brought). The old computer_20251124 tool is refused on the Claude API and Google Cloud, though Bedrock still takes it. Thinking blocks are bound to the model and the conversation, so moving a chat from Sonnet 5.5 to any other model drops its reasoning, and editing an earlier turn can trigger a 400 on accounts created from 31 August 2026. And the advisor tool won't accept Opus 4.8, Opus 4.7 or Sonnet 5 as advisors any more.

Then there's the quiet one. Longer notes between tool calls now come back as thinking blocks, and at the default display: "omitted" their text is empty. Your agent UI just stops talking between steps. No error. With between_tools the text comes back, which honestly makes it the setting we'd start from for chat style agents.

Should you move now?

For anyone already on Sonnet 5, yes, once the five checks pass. Same price, same tokenizer (so the same token counts), a 512 token minimum for caching instead of 1,024, and per-message effort plus mid-conversation system messages, which Sonnet 5 didn't have. Don't carry your effort setting over, though. Anthropic says the levels are recalibrated, and suggests starting agentic coding at medium rather than the API default of high.

For anyone on Opus 5.5, it's less obvious. On the published numbers you're paying double for a couple of points, and on Terminal-Bench for nothing. We'd run a side by side on real tasks before downgrading, because one early independent test didn't go smoothly: Simon Willison's max effort run burned all 128K output tokens ($1.28) without finishing, while xhigh did the job in 41 seconds for under six cents. He also saw the extended thinking bug he'd hit on Opus 5.5. Sonnet 5.5 is now the default on the free claude.ai tier, per Willison. Haiku 5.5 is due "in the coming weeks", so the bottom of the range isn't settled either.

A quick check that your key sees the model:

Linux
curl -s https://api.anthropic.com/v1/models/claude-sonnet-5-5 -H "x-api-key: $ANTHROPIC_API_KEY" -H "anthropic-version: 2023-06-01"

Sources

Price, cache rates, context, output limits, cutoff, retirement date and model IDs come from Anthropic's Claude Sonnet 5.5 model page. The five breaking changes, between_tools rules, recalibrated effort and the 512 token cache minimum are from What's new in Claude Sonnet 5.5. Benchmarks, footnotes, the speed and cost claims and the Haiku 5.5 timing are from the Sonnet 5.5 launch page. Copilot availability is from the GitHub changelog. The max and xhigh test runs and the free tier default are from Simon Willison. The figures are ours, built from those pages.

Frequently asked questions

How much does Claude Sonnet 5.5 cost?

$2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. Cache reads are $0.20, five minute cache writes $2.50, one hour writes $4, and batch requests are half price. Those are Anthropic's list prices on 29 September 2026.

Can I still turn thinking off on Sonnet 5.5?

Not with disabled, which now returns a 400. Send thinking: {"type": "between_tools"} instead, at high effort or below. It removes up-front thinking but still returns short progress notes between tool calls.

Is Sonnet 5.5 as good as Opus 5.5?

Close, on Anthropic's own table. It trails by about two points on CursorBench 4.0 and OSWorld 2.1, ties within two Elo on GDPval-AA, and leads on Terminal-Bench 4.0. Those are vendor figures, so we'd test both on your own workload.

Is Claude Sonnet 5.5 in GitHub Copilot?

Yes, rolling out gradually from 28 September to Copilot Pro, Pro+, Max, Business and Enterprise, billed at the provider's list price under usage-based billing, per GitHub's changelog.

Tags: AI pricinganthropicclaudellmnews
Share198Tweet124
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card: Reflection AI announced Beam, a 501B parameter open-weight mixture of experts model, with weights promised later in October 2026.

Reflection Beam is a 501B open model with no weights yet

6 October 2026
Answer card: Aleph Alpha released Kolibri-1 on 3 October 2026 as an Apache 2.0 mixture of experts model with 78.1 billion total and 3.46 billion active parameters, served at 1 million tokens of context but trained on sequences of up to 256 thousand tokens.

Kolibri-1 serves 1M tokens but was trained to 256K

4 October 2026
Answer card on Gemini 4 Argon pricing and gated access, announced by Google on 30 September 2026.

Gemini 4 Argon costs $2 and $10 now, $4 and $20 later

2 October 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.