If your Sonnet 5 code sends thinking: {"type": "disabled"}, Claude Sonnet 5.5 will hand it back as a 400. That's the first thing we'd check. The model itself went live on 28 September as claude-sonnet-5-5, at the same $2 per million input tokens and $10 per million output as Sonnet 5, and on Anthropic's own table it sits within two or three points of Opus 5.5, a model that costs twice as much. Cheap and close to the top. Also a bit stricter about how you call it.
The short answer
Claude Sonnet 5.5 shipped 28 September 2026 at $2 and $10 per million tokens, cache reads at $0.20, with a 1M context window and 128K output. It scores 70.6% on Terminal-Bench 4.0, ahead of Opus 5.5, and trails it by a couple of points elsewhere. Before switching, replace disabled thinking with between_tools and drop forced tool_choice.
The cheap model is now close to the expensive one
Nothing moved on the price card. Input is $2, output $10, five minute cache writes $2.50, one hour writes $4, cache reads $0.20, and batch halves all of it. The window is 1M tokens, max output is 128K on the synchronous API, and batch can go to 300K with the output-300k-2026-03-24 beta header. Knowledge cutoff is June 2026. Retirement is "not sooner than" 28 September 2027.
What moved is the gap to Opus 5.5. On Anthropic's published table Sonnet 5.5 gets 55.5% on CursorBench 4.0 against 57.8%, 80.1% on OSWorld 2.1 against 81.8%, and 1844 Elo on GDPval-AA against 1846. Two Elo points. On Terminal-Bench 4.0 it actually leads, 70.6% against 66.4% for Opus 5.5 at xhigh effort. For a model at half the per token price, that's a strange place for the flagship to be.
Here's where I'd slow down. The same table gives Sonnet 5 just 10.3% on Terminal-Bench 4.0, which would make this a sevenfold jump in one point release. Maybe it is. I'd bet more on a harness or setting that Sonnet 5 handled badly, and the page doesn't explain it. These are vendor numbers, and Artificial Analysis flagged that it tested a pre-release build with a structured output bug. There's also a footnote we liked for its honesty: on FrontierCode, max effort scored lower than xhigh (46.2% against 52.1%), because more effort sometimes sent the model into code review loops that ended in timeouts or out of scope edits.
The efficiency claim is the one that matters for a bill. Anthropic says output is "30%+ faster" and a task costs "up to 30% less", mostly through fewer tokens and tool calls. GitHub's changelog says the same thing from its side: it matched Sonnet 5 on coding tasks with "significantly fewer steps". Up to. We'd still measure it on our own traces before believing the invoice.
between_tools, and the other four breaks
If you read our Opus 5.5 piece last week, most of this will look familiar. Sonnet gets the same treatment with one twist. Thinking still has an off switch, it's just renamed and narrower. You send thinking: {"type": "between_tools"}, which kills the up-front reasoning but keeps the short progress notes between tool calls. It only works at low, medium or high effort. At xhigh or max it's a 400, it won't take display or budget_tokens, and effort can't change mid-conversation while it's on. Without tools, you get plain text back, like disabled did.
The other four. Forced tool use is gone: tool_choice of any or a named tool returns a 400, even on token counting, so structured extraction jobs need auto plus strict tools or structured outputs (the same break Fable 5.1 brought). The old computer_20251124 tool is refused on the Claude API and Google Cloud, though Bedrock still takes it. Thinking blocks are bound to the model and the conversation, so moving a chat from Sonnet 5.5 to any other model drops its reasoning, and editing an earlier turn can trigger a 400 on accounts created from 31 August 2026. And the advisor tool won't accept Opus 4.8, Opus 4.7 or Sonnet 5 as advisors any more.
Then there's the quiet one. Longer notes between tool calls now come back as thinking blocks, and at the default display: "omitted" their text is empty. Your agent UI just stops talking between steps. No error. With between_tools the text comes back, which honestly makes it the setting we'd start from for chat style agents.
Should you move now?
For anyone already on Sonnet 5, yes, once the five checks pass. Same price, same tokenizer (so the same token counts), a 512 token minimum for caching instead of 1,024, and per-message effort plus mid-conversation system messages, which Sonnet 5 didn't have. Don't carry your effort setting over, though. Anthropic says the levels are recalibrated, and suggests starting agentic coding at medium rather than the API default of high.
For anyone on Opus 5.5, it's less obvious. On the published numbers you're paying double for a couple of points, and on Terminal-Bench for nothing. We'd run a side by side on real tasks before downgrading, because one early independent test didn't go smoothly: Simon Willison's max effort run burned all 128K output tokens ($1.28) without finishing, while xhigh did the job in 41 seconds for under six cents. He also saw the extended thinking bug he'd hit on Opus 5.5. Sonnet 5.5 is now the default on the free claude.ai tier, per Willison. Haiku 5.5 is due "in the coming weeks", so the bottom of the range isn't settled either.
A quick check that your key sees the model:
curl -s https://api.anthropic.com/v1/models/claude-sonnet-5-5 -H "x-api-key: $ANTHROPIC_API_KEY" -H "anthropic-version: 2023-06-01"
Sources
Price, cache rates, context, output limits, cutoff, retirement date and model IDs come from Anthropic's Claude Sonnet 5.5 model page. The five breaking changes, between_tools rules, recalibrated effort and the 512 token cache minimum are from What's new in Claude Sonnet 5.5. Benchmarks, footnotes, the speed and cost claims and the Haiku 5.5 timing are from the Sonnet 5.5 launch page. Copilot availability is from the GitHub changelog. The max and xhigh test runs and the free tier default are from Simon Willison. The figures are ours, built from those pages.
Frequently asked questions
How much does Claude Sonnet 5.5 cost?
$2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. Cache reads are $0.20, five minute cache writes $2.50, one hour writes $4, and batch requests are half price. Those are Anthropic's list prices on 29 September 2026.
Can I still turn thinking off on Sonnet 5.5?
Not with disabled, which now returns a 400. Send thinking: {"type": "between_tools"} instead, at high effort or below. It removes up-front thinking but still returns short progress notes between tool calls.
Is Sonnet 5.5 as good as Opus 5.5?
Close, on Anthropic's own table. It trails by about two points on CursorBench 4.0 and OSWorld 2.1, ties within two Elo on GDPval-AA, and leads on Terminal-Bench 4.0. Those are vendor figures, so we'd test both on your own workload.
Is Claude Sonnet 5.5 in GitHub Copilot?
Yes, rolling out gradually from 28 September to Copilot Pro, Pro+, Max, Business and Enterprise, billed at the provider's list price under usage-based billing, per GitHub's changelog.






















