• Latest
  • Trending
  • All
Answer card: Fable 5 costs five times Sonnet 5 during the intro window and buys a 17-point SWE-bench Pro lead; Sonnet 5 remains the right daily driver.

Claude Sonnet 5 vs Fable 5: is the top model worth 5x?

2 July 2026
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Tuesday, September 22, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Claude Sonnet 5 vs Fable 5: is the top model worth 5x?

by stephane
2 July 2026
in Dev
0
Answer card: Fable 5 costs five times Sonnet 5 during the intro window and buys a 17-point SWE-bench Pro lead; Sonnet 5 remains the right daily driver.
493
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Anthropic pulled off a strange double feature. On June 30 it launched Sonnet 5 at $2 in and $10 out per million tokens, and one day later Fable 5 came back from its export-control freeze at $10 and $50. Same family, same week, same 1 million token context, and one costs exactly five times the other. So every Claude Code user now faces the same fork: the cheap default that wins agentic benchmarks, or the flagship that crushes deep coding and runs for days. The honest answer is not one model. Sonnet 5 is enough for most of what you do, Fable 5 is untouchable on the hardest tenth, and the trick is knowing which tenth. Here is the gap, priced and mapped.

The short answer

Same week, same family, five times the price. Sonnet 5 ($2/$10 intro) is the default and wins the published agentic numbers; Fable 5 ($10/$50) crushes deep coding with a 17-point SWE-bench Pro lead and runs multi-day tasks nothing else sustains. Daily driver: Sonnet 5. The hardest tenth of your work: Fable 5, deliberately, at a chosen effort level.

5xthe price gap, intro window
80.3 vs 63.2SWE-bench Pro
1Mcontext on both
Answer card: Fable 5 costs five times Sonnet 5 and buys a 17-point SWE-bench Pro lead; Sonnet 5 stays the daily driver, Fable 5 takes the hardest tenth.
One family, one week apart, a 5x price gap. The split that matters is by task, not by badge.

One week, two models, one fork

The timing was almost comic. June 30: Sonnet 5 launches, becomes the default in Claude Code and on the free plans, and undercuts everything with a $2/$10 intro price that runs through August. July 1: Fable 5, Anthropic’s Mythos-class flagship, returns from nineteen days under a US export-control order at $10/$50. Two brand-new options in the same picker, twenty-four hours apart, priced exactly five times apart.

On paper they look closer than the price says. Both carry a 1 million token context window and 128K output. Both run adaptive thinking by default, both take the same effort dial up to xhigh. The spec sheet will not help you choose. The benchmarks and the bill will.

What the 5x actually buys

Start with the number that justifies Fable 5’s existence. On SWE-bench Pro, the hard multi-file repository benchmark, Fable 5 posts 80.3 against Sonnet 5’s 63.2. Seventeen points. For scale, the gap between Sonnet 5 and Opus 4.8 on the same test is six points, and we called that a real lead worth paying for on the hardest work. Fable 5 nearly triples that lead over the mid-tier. On SWE-bench Verified it is the first model past 90, at 95 against Sonnet 5’s 85.2. And on the long-horizon evals, Anthropic reports it breaking 90 percent on complex multi-step analytical work and topping Cognition’s FrontierBench, the territory of tasks that run for hours or days without a human nudging them along.

That last part is the honest heart of the pitch. Fable 5 is not a slightly smarter Sonnet. It is built for a different shape of work: the migration that touches four hundred files, the agent session that survives overnight, the research question that needs two hundred tool calls to answer properly. Sonnet 5 does not lose that work gracefully. It loses the thread, and you re-prompt, and the savings evaporate into your afternoon.

Grouped bar chart: SWE-bench Pro, Sonnet 5 63.2 vs Fable 5 80.3; SWE-bench Verified, Sonnet 5 85.2 vs Fable 5 95.
The two shared coding benchmarks. The 17-point Pro gap is the whole argument for the flagship.

What Sonnet 5 keeps anyway

Now the other side, because it is stronger than the price gap implies. Sonnet 5 holds the highest published Terminal-Bench score in the family at 80.4, ahead of Opus 4.8’s 74.6, and Fable 5’s number has simply not been published, an absence worth noticing in a launch that quoted every figure it liked. On knowledge work Sonnet 5 edges Opus on GDPval. These are the benchmarks closest to what an everyday Claude Code session actually does: drive a terminal, edit files, run tests, iterate.

There is also a verification asymmetry to keep in mind. Sonnet 5’s numbers have been picked over publicly for days; Fable 5’s headline figures are Anthropic-reported, with independent reproduction still thin and at least one benchmark conspicuously missing. None of that means the flagship is oversold. It means the exact size of the gap is provisional, and the direction is not.

Bar chart pricing the same 20k-in, 20k-out turn: Sonnet 5 intro $0.24, Sonnet 5 standard $0.36, Fable 5 $1.20.
The same turn, priced at each rate card. Computed at published rates; Fable 5 at xhigh effort can triple its own bar.

The bill, on the same task

Price the same turn on both rate cards and the gap stops being abstract. Take an agentic turn of 20,000 tokens in and 20,000 out. On Sonnet 5’s intro pricing it costs $0.24. On the standard pricing that starts in September, $0.36. On Fable 5, $1.20. A hundred-turn session: $24 versus $120, same work, and that assumes Fable 5 behaves. It does not always, in the sense that it always thinks, bills the thinking as output, and at xhigh effort the same turn runs closer to $3.20. Push the flagship’s dial and the daily gap is not 5x. It is 13x.

Scale it to projects and the fork gets vivid. The same 1,500-word article, five turns: about $1.20 on Sonnet 5, $6 on Fable 5. A small app, eighty turns: roughly $19 versus $96. A heavy overnight migration, three hundred turns: about $72 versus $360 and up, except that this last one is exactly the job where Sonnet 5 can lose the thread partway through, and one lost thread erases the whole saving. That is the entire decision, compressed: the cheaper model, until the task is the kind that fails expensive.

On the subscription side the asymmetry is gentler but real: Sonnet 5 is the free default everywhere, while Fable 5 rides paid plans at up to half your weekly limits through July 7 and moves to usage credits after. Either way, the flagship is metered in a way the default is not, which is exactly how Anthropic is telling you to use it.

The split that works

So run the fleet like this. Sonnet 5 is the daily driver, and since it is already the default you have to do precisely nothing: everyday coding, agent runs, drafts, reviews, the 90 percent of work where the mid-tier’s answer and the flagship’s answer are indistinguishable except on the invoice. Fable 5 is the specialist, summoned deliberately: the multi-file migration, the overnight agent, the research task with real depth, anything where Sonnet 5 already fell short once. Summon it at high effort, not xhigh, until a run proves it needs more.

And if a task sits in between, there is a middle you already know about: Opus 4.8 at $5/$25 still splits the difference on both capability and cost. The family suddenly has a real ladder again. The skill is climbing it only as far as the task demands, because the view from the top is expensive, and most days you do not need it.

Sources: Anthropic’s Sonnet 5 and Fable 5 announcements and the Claude Platform docs; benchmark tables collated by MarkTechPost and MorphLLM, June and July 2026. Fable 5 figures are largely Anthropic-reported; dollar figures are computed from published rates at equal token volumes.

Frequently asked questions

Is Claude Fable 5 better than Sonnet 5?

On raw capability, clearly: Fable 5 scores 80.3 on SWE-bench Pro against Sonnet 5's 63.2, and 95 percent on SWE-bench Verified against 85.2. It is built for long-horizon, multi-day agentic work no mid-tier model sustains. But it costs five times as much during Sonnet 5's intro pricing, so better is not the same as worth it for everyday tasks.

How much more does Fable 5 cost than Sonnet 5?

Exactly 5x during Sonnet 5's intro window: $10/$50 per million tokens versus $2/$10, through August 31, 2026. After that Sonnet 5 moves to $3/$15, making Fable 5 about 3.3x. And because Fable 5 always thinks and bills thinking as output, its real-world gap on long tasks is often wider than the rate card suggests.

Which model should I use in Claude Code?

Sonnet 5 as the default, which it already is. It wins the published terminal and agentic numbers in the family and costs a fraction of the flagship. Switch to Fable 5 for the hardest work: large multi-file migrations, tasks that run for hours or days, and problems where Sonnet 5 measurably fell short. On paid plans, Fable 5 is included up to 50 percent of weekly limits through July 7, then moves to usage credits.

Do Sonnet 5 and Fable 5 have the same context window?

Yes, both offer 1 million tokens of context with 128K max output, and both run adaptive thinking by default with the same effort levels, including xhigh. The difference is not the specs sheet, it is how far each can push a hard problem and what a token of it costs.

Are Fable 5's benchmark numbers verified?

Mostly Anthropic-reported so far, and worth reading that way. The SWE-bench Pro 80.3 and SWE-bench Verified 95 figures come from Anthropic; independent reproductions are still limited, and Fable 5's Terminal-Bench score has not been published at all. The direction of the gap is not in doubt, but treat exact deltas as provisional.

Tags: aianthropicarticlebenchmarksclaudellm
Share197Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.