• Latest
  • Trending
  • All
Answer card: Claude Fable 5 is back worldwide; its per-token rate is fixed but the effort dial moves the real bill by about 7x between low and xhigh.

Claude Fable 5 is back: what each effort level costs

3 September 2026
Answer card for Claude Opus 5.5: 4 dollars in and 20 dollars out per million tokens, down from 5 and 25, with cache reads at 20 cents.

Claude Opus 5.5 drops to $4 and $20, and breaks four Opus 5 habits

23 September 2026
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Wednesday, September 23, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Claude Fable 5 is back: what each effort level costs

by stephane
3 September 2026
in Dev
0
Answer card: Claude Fable 5 is back worldwide; its per-token rate is fixed but the effort dial moves the real bill by about 7x between low and xhigh.
497
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Claude Fable 5 lasted three days. Launched June 9, suspended June 12 under a US export control order, and as of July 1 it’s back worldwide, on Claude.ai, Claude Code and the API. So the obvious question returns with it: what does the most capable model on the market actually cost to run? The sticker says $10 in and $50 out per million tokens, and the sticker is misleading. Fable 5 thinks on every request, and the effort dial (low, medium, high, xhigh, max) changes how many tokens it burns to answer you. Same task, same rates, and the bill moves by 7x between settings. Nobody prices that out level by level, so we did. Here’s the story of the ban, the math of the dial, and which effort each kind of work deserves.

The short answer

Fable 5 is back worldwide after the US lifted its export controls, with a new safety classifier as the price of return. The per-token rate never changes: what changes is how many tokens each effort level burns. Same agentic turn, roughly $0.45 on low, $1.20 on high, $3.20 on xhigh. Default to high, drop to low for the mechanical stuff, and save xhigh for work that measurably needs it.

19 daysbanned, of its first 22
$10 / $50per 1M tokens, fixed
~7xthe bill, low vs xhigh
Answer card: Fable 5 is back worldwide; the rate is fixed at $10 in and $50 out per million tokens, and the effort dial moves the real bill by about 7x.
The rate card isn’t the price. The effort dial is.

Three days on the market, nineteen in a vault

The timeline reads like fiction. Anthropic shipped Claude Fable 5 on June 9, its first Mythos-class model for the general public, state of the art on coding, knowledge work and computer use. On June 12 the US government applied export controls and global access went dark. The trigger: Amazon researchers had found a prompting technique that walked past the model’s safeguards and got it hunting software vulnerabilities, in at least one case producing code that demonstrated a working exploit. For a model this capable, that report was enough.

The controls were lifted on June 30 and Fable 5 came back worldwide on July 1, on Claude.ai, Claude Code, Claude Cowork and the API, with cloud platforms following. The return came with conditions worth knowing about: deeper collaboration with the US government, including pre-release access, and a new safety classifier that blocks the reported technique in over 99 percent of cases. The honest footnote is that the classifier also raises false positives on routine security and coding work, so if a benign request gets declined, that’s why. The API ships an opt-in fallback that re-serves declined requests on Opus 4.8 inside the same call, which is worth turning on if you automate anything.

Availability on the plans is time-boxed generosity: Pro, Max, Team and select Enterprise get Fable 5 included for up to half their weekly usage limits through July 7, after which it moves to usage credits. The API is straight per-token pricing. Which brings us to the actual subject.

The rate is fixed. The bill isn’t.

Fable 5 costs $10 per million input tokens and $50 per million output, double Opus 4.8’s $5 and $25. Most coverage stops there, and stopping there’s how people end up shocked by an invoice, because Fable 5 has a property that makes the rate card almost decorative: thinking is always on. You can’t turn it off. Every request reasons before it answers, the depth of that reasoning is steered by one parameter, effort, and every thinking token bills as output even when you never see it.

Effort takes five values: low, medium, high (the default), xhigh and max, with xhigh reserved for the top models (Fable 5, Mythos 5, Opus 4.8 and 4.7). It’s not a price multiplier. It’s a behavior multiplier. At low effort the model skips preamble, combines tool calls and answers tersely. At xhigh it plans out loud, explores, double-checks, and spends tokens like a researcher on a deadline. Rates fixed, volume anything but.

Bar chart: the same agentic turn on Fable 5 costs about $0.45 at low effort (5k output tokens), $1.20 at high (20k), and $3.20 at xhigh (60k), computed at $10 and $50 per million tokens.
Same task, same rate card, 7x the bill. Computed from published per-effort token volumes, stated on the figure.

The same task, priced at every effort

Numbers make it concrete. Take a typical agentic turn with 20,000 tokens of input, the kind Claude Code fires constantly, and use the per-effort output volumes documented from real usage: about 5,000 output tokens on low, 20,000 on high, 60,000 on xhigh. Price them at $10 and $50 per million and you get, per turn:

The input costs $0.20 every time. On low, 5,000 output tokens add $0.25: about $0.45 a turn. On high, 20,000 add $1.00: about $1.20. On xhigh, 60,000 add $3.00: about $3.20. Medium lands between low and high, and max has no published ceiling at all, by design: Anthropic describes it as no constraints on token spending. Run a hundred-turn agent session and the gap stops being pocket change: about $45 on low, $120 on high, $320 on xhigh, for the same hundred turns.

Two levers cut all of these regardless of effort. Prompt caching takes up to 90 percent off repeated input, which in agent loops is most of it, and batch mode halves anything that can wait. But nothing you do to the input side changes the core fact: on a model that always thinks, the effort dial is the price tag, and it’s set per request, by you.

Price it like a project, not a turn

Per-turn numbers stay abstract until you multiply them by a real job. So take four familiar projects and price them end to end, same assumptions as above: about 20,000 tokens of input per turn, the published per-effort output volumes. The turn counts are honest estimates from real Claude Code sessions; yours will vary, the ratios won’t.

Writing a 1,500-word article, five turns of draft and revisions: about $2 on low, $6 on high, $16 on xhigh. Building a landing page, forty turns of scaffold, styling and fixes: $18, $48, $128. A small app, say a subnet calculator with a web UI, eighty turns: $36, $96, $256. A heavy build or migration, hundreds of files, an agent running overnight, three hundred turns and realistically the reason you rented Fable 5 at all: $360 at high, and pushing $960 at xhigh.

Table pricing four projects end to end on Fable 5 at each effort level, with the recommended setting boxed per project: low for an article, high for a landing page and a small app, xhigh for a heavy migration.
Four real projects, priced at every effort. The boxed price is the setting to start with.

And here’s the pick, per project, so nobody has to guess. The article: low. Fable 5’s low effort already writes at the level of yesterday’s flagships, and prose doesn’t need sixty thousand tokens of deliberation. The landing page and the small app: high, the default, which is exactly what agentic coding was tuned for. The heavy build or migration: xhigh, the one job whose failure costs more than its tokens. Which makes the realistic bills $2, $48, $96 and $960, not four times the worst case.

Two things jump out of that table. First, the dial matters more than the project: an article at xhigh costs almost what a landing page costs at low. Second, at these volumes prompt caching stops being a nice-to-have, because in a real session most of that per-turn input is cache reads at a tenth of the rate. The output side, the thinking, is the part no cache can save you from.

Which effort for which work

Anthropic publishes recommendations, and they match what the math suggests.

Low is for the mechanical layer: subagents doing one bounded thing, classification, extraction, quick lookups, latency-sensitive paths where a fast terse answer beats a considered one. A striking detail from the docs: Fable 5’s lower efforts often outperform the xhigh of previous-generation models, so low here’s not the discount bin. It’s yesterday’s flagship at a sixth of the spend.

Medium fits everyday drafting, summaries and Q&A, work where you want some deliberation but the problem isn’t deep. High, the default, is the right home for daily coding and standard agent runs, and if you never touch the dial this is what you’re already paying. xhigh is for capability-sensitive, long-horizon work: the multi-day migration, the refactor that has to hold a whole system in its head, deep research where a missed connection costs more than the tokens. And max is for frontier problems only. Anthropic’s own framing is that it adds significant cost for relatively small gains, which is a vendor telling you not to buy its most expensive setting unless you must.

Checklist matching effort levels to use cases: low for subagents and classification, medium for everyday drafting, high as the daily default, xhigh for long-horizon capability-sensitive work, max for frontier problems only, and never blanket xhigh out of habit.
The dial, mapped to the work. The one anti-pattern: xhigh everywhere, out of habit.

The habit that costs the most

If one thing on this page saves you money, let it be this: the expensive mistake isn’t picking xhigh for a hard problem. It’s leaving xhigh on for everything because it worked once. Blanket maximum effort turns a $45 agent session into a $320 one and buys you longer preambles on tasks a terse answer would have served. The discipline that works is boring: default to high, drop to low wherever the task is mechanical, and promote a task to xhigh only when a run at high measurably fell short. Check the usage field in the response, thinking_tokens sits right there, and treat any task that didn’t need its thinking budget as a candidate for demotion.

Fable 5 coming back is genuinely good news: it’s the most capable model generally available, and for the hardest work nothing else currently touches it. Just remember that its real price isn’t on the rate card. It’s on a dial, the dial defaults to sensible, and every notch above sensible should be a decision, not a habit. For whether this much model is worth five times the new mid-tier, our Sonnet 5 vs Fable 5 head-to-head prices that exact fork, and Sonnet 5 vs Opus 4.8 covers the middle of the ladder.

Sources: Anthropic’s Fable 5 redeployment announcement and Claude Platform docs; export-control timeline via VentureBeat and Forbes; per-effort token volumes collated by Developers Digest, July 2026. Dollar figures on this page are computed from those volumes at the published $10/$50 rates; your tasks will vary.

Frequently asked questions

Why was Claude Fable 5 suspended?

A US export control order landed on June 12, three days after launch, after Amazon researchers reported a prompting technique that bypassed its safeguards and got it to find software vulnerabilities, in one case producing working exploit code. Anthropic suspended global access, added a classifier that blocks the reported technique in over 99 percent of cases, and the controls were lifted on June 30.

How much does Claude Fable 5 cost?

The API rate is $10 per million input tokens and $50 per million output, double Opus 4.8. But the real bill depends on the effort setting, because thinking tokens are billed as output even when hidden: the same agentic turn that costs about $0.45 on low effort runs about $1.20 on high and about $3.20 on xhigh. On Claude plans, Fable 5 counts against usage limits, included up to 50 percent of weekly limits through July 7, then via usage credits.

What are the effort levels on Claude Fable 5?

Five: low, medium, high, xhigh and max. High is the default. xhigh exists only on Fable 5, Mythos 5, Opus 4.8 and Opus 4.7. Effort isn’t a price multiplier; it steers how many tokens the model spends thinking, calling tools and explaining, which is what actually moves the bill.

Which effort level should I use?

Anthropic's own guidance: start at high, the default, and reserve xhigh for capability-sensitive, long-horizon work. Use low for subagents, classification and quick lookups. Fable 5's lower efforts often beat the xhigh of previous models, so resist running everything at the ceiling out of habit. Max is for frontier problems and adds a lot of cost for small gains.

Does Claude Fable 5 always think?

Yes. Thinking is always on and adaptive; you can’t disable it, and explicit thinking budgets are rejected. You control depth through the effort parameter instead, and the thinking tokens bill as output even when the summary is hidden. That’s exactly why effort, not the rate card, is the real cost lever.

Tags: aianthropicarticleclaudellmpricing
Share199Tweet124
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card for Claude Opus 5.5: 4 dollars in and 20 dollars out per million tokens, down from 5 and 25, with cache reads at 20 cents.

Claude Opus 5.5 drops to $4 and $20, and breaks four Opus 5 habits

23 September 2026
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.