• Latest
  • Trending
  • All
Answer card: DeepSeek replaced the weights behind deepseek-v4-pro with DeepSeek-V4-Pro-0813 on 12 August 2026, keeping the same endpoint and the same price of 0.435 dollars per million input tokens and 0.87 per million output, with no change log entry and no 0813 weights on Hugging Face.

DeepSeek V4 Pro 0813 is live, at 3x the Flash price

12 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Tuesday, September 22, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

DeepSeek V4 Pro 0813 is live, at 3x the Flash price

by stephane
12 September 2026
in Dev
0
Answer card: DeepSeek replaced the weights behind deepseek-v4-pro with DeepSeek-V4-Pro-0813 on 12 August 2026, keeping the same endpoint and the same price of 0.435 dollars per million input tokens and 0.87 per million output, with no change log entry and no 0813 weights on Hugging Face.
492
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Nothing in your code has to change, again. On 12 August DeepSeek pointed deepseek-v4-pro at a new build called DeepSeek-V4-Pro-0813, same endpoint, same 0.435 dollars in and 0.87 out, new weights underneath. We found out the way everyone else did, by noticing a table on the pricing page had changed, because there's no change log entry and no blog post. Flash got the same treatment on 31 July, and at least that one came with a note. The swap itself is routine by now. What's worth your attention is what it does to the choice between the two DeepSeek models, because for the past two weeks the cheap one had been beating the expensive one, and that's over.

The short answer

DeepSeek-V4-Pro-0813 now answers calls to deepseek-v4-pro. The endpoint, the price and the 1M context are unchanged, and the swap arrived with no change log entry. The 0813 weights are not on Hugging Face, which breaks the pattern every previous V4 release set. Pro is faster than Flash on DeepSeek’s own agent benchmarks again, at a little over three times the cost per token.

12 Augweights swapped in place
3.1xPro token price against Flash
500concurrent requests, Flash gets 2500
Answer card: DeepSeek replaced the weights behind deepseek-v4-pro with DeepSeek-V4-Pro-0813 on 12 August 2026, at the same endpoint and the same price of 0.435 dollars per million input and 0.87 per million output, with no change log entry and no 0813 weights on Hugging Face.
A version string in a docs table is the whole announcement so far.

What actually changed

We compared the models and pricing page against its own archived copy. On 9 August the model version column read DeepSeek-V4-Pro. Today it reads DeepSeek-V4-Pro-0813. That’s it. That’s the announcement.

One other row moved with it, and it’s the one developers were waiting on. The Responses API was flash-only until now, with a footnote promising deepseek-v4-pro support in early August. That footnote is gone and the row is a tick. So if you were holding a Responses API migration because your hard calls go to Pro, you can stop holding.

Everything else on the page held still. Same 1M context, same 384K ceiling on output, same prices to four decimal places, same concurrency limits. The change log still ends at 31 July, which was the Flash update. Nothing about 0813.

Checklist comparing what changed on the DeepSeek pricing page between 9 and 12 August 2026, namely the model version string and Responses API support for deepseek-v4-pro, against what did not change, namely the prices, the 500 request concurrency limit, the absence of a change log entry and the absence of 0813 weights on Hugging Face.
Two ticks moved. The right-hand column is the part worth arguing about.

The weights broke the pattern

Here’s the bit we didn’t expect. Every V4 release so far shipped with weights. The April preview put both Pro and Flash on Hugging Face under MIT on day one. Flash 0731 did the same, ungated, with a vLLM recipe in the model card.

There is no DeepSeek-V4-Pro-0813 repository. We checked the Hugging Face API directly rather than trusting a search box, and the newest thing on the deepseek-ai account is still Flash 0731, uploaded 1 August. So the weights you can download today are the April preview. Not what the API is serving.

Maybe they land next week and this paragraph ages badly. I’d guess they do, honestly, because opening the weights is most of DeepSeek’s leverage and they’ve never skipped it. But if you self-host V4-Pro and you’ve been telling people it matches the hosted model, that stopped being true on 12 August, and nobody sent you a note.

Pro takes the lead back from Flash

Two weeks ago the awkward fact about DeepSeek’s lineup was that the small model beat the big one. Flash 0731 outscored the V4-Pro preview on every agent benchmark DeepSeek published, at a third of the price, which made Pro difficult to recommend for anything.

The numbers going round for 0813 fix that. Terminal Bench 2.1 at 87.9 where the preview managed 72.1. DeepSWE at 62.7 from 12.8. AutomationBench 31.8 from 12.8, DSBench-Hard 67.2 from 31.1. Those come from a DeepSeek chart circulating on X and reported by trade press, and we could not find them on any DeepSeek page. Read them as DeepSeek’s claim about DeepSeek, because that’s what they are.

Bar chart of Terminal Bench 2.1 scores reported by DeepSeek for three of its own builds: V4-Pro preview at 72.1 in April, V4-Flash-0731 at 82.7 in July, and V4-Pro-0813 at 87.9 in August.
Ahead again. By 5.2 points, for a little over three times the price per token.

Put the price next to it and the decision gets simpler, not harder. Pro is $0.435 per million input and $0.87 output. Flash is $0.14 and $0.28. That ratio is exactly 3.107 on both sides, which is a suspiciously tidy number and probably deliberate. You’re paying triple for about five points on a benchmark DeepSeek ran itself.

The concurrency limit is the part people miss. Pro allows 500 concurrent requests, Flash allows 2500. If you’re running a batch job or a fan-out agent, that ceiling will hurt you long before the bill does.

The price warning is still sitting there

Same page, footnote one, and it hasn’t moved since early August:

We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.

Worth knowing what that footnote replaced. Earlier in August the same slot held a concrete plan: peak and off-peak billing, 2x the regular rate between 09:00 and 12:00 and again 14:00 to 18:00 Beijing time. That’s gone, swapped for a blanket warning with no number and no date. Whether that’s better or worse depends entirely on how much of your traffic was going to land in those windows.

Either way, don’t model a year of spend on today’s rates. DeepSeek has told you in writing that they’re temporary.

Should you switch anything

If your Pro calls are already going to deepseek-v4-pro, you’ve switched. That’s the nature of an in-place swap, and it’s the same trap we wrote about when the legacy aliases retired in July. Re-run whatever evals you trust. The results you have on file describe the April preview.

If you’d moved hard work to Flash because Pro looked bad value, it’s worth another look, though not automatically. Five benchmark points is real on tasks that were failing outright and invisible on tasks that were already passing. Try it on the calls that actually hurt.

And if you self-host, sit tight. The build worth downloading isn’t downloadable yet.

What changed since

12 September 2026. Three things above have moved. The weights first, because we were wrong to worry: DeepSeek put a DeepSeek-V4-Pro-0813 repository on Hugging Face on 13 August, the day after this piece went up, so the build the API serves is downloadable after all. Then the price. The warning we quoted turned into peak and off-peak billing on 16 August, and deepseek-v4-pro now costs $0.66 in and $1.98 out off-peak, $1.32 and $3.96 at peak, against the $0.435 and $0.87 shown here. And the model’s future. When DeepSeek shipped DeepSeek-V4.1-Flash on 10 September, its launch post said every deepseek-v4-pro request would be routed to the new Flash at Flash rates from 14 September, until a V4.1-Pro arrived. A day later the pricing page grew a footnote: in response to user demand, V4 Pro stays on the API after 14 September with billing unchanged. The details are in our V4.1-Flash write-up.

Sources

Model version, pricing, concurrency limits and the price increase notice are from DeepSeek’s own models and pricing page, compared against the archived copy of 9 August 2026 for the diff. The GA description, the endpoint creation time of 15:42 UTC on 12 August and the context and output limits are from the OpenRouter listing and its public models API. Weight availability was checked against the deepseek-ai account on Hugging Face. The 0813 benchmark figures are DeepSeek-reported, circulated via a widely shared post on X and covered by Wccftech, and we have not reproduced them. The preview and Flash 0731 scores are from DeepSeek’s published table of 31 July.

Frequently asked questions

What is DeepSeek-V4-Pro-0813?

It is the build now sitting behind the deepseek-v4-pro model name in DeepSeek's API. The models and pricing page listed the version as plain DeepSeek-V4-Pro as recently as 9 August and shows DeepSeek-V4-Pro-0813 on 12 August. OpenRouter listed a matching deepseek-v4-pro-0813 endpoint at 15:42 UTC on 12 August and describes it as the GA release of DeepSeek V4 Pro, which is the closest thing to an announcement anyone has.

Do I need to change my API calls?

No. The model name is still deepseek-v4-pro, the base URLs are unchanged, and the 1M context and 384K maximum output are the same. That is exactly why it is worth a note in your own log: anything already pointed at that name picked up new weights without an opt-in. If you keep an eval suite from before 12 August, it describes a model you are no longer talking to.

How much does DeepSeek V4 Pro cost now?

Per million tokens, DeepSeek publishes $0.003625 on a cache hit, $0.435 on a cache miss and $0.87 for output. Those numbers did not move with the 0813 swap. They are 3.1x the deepseek-v4-flash rates of $0.14 and $0.28. The same page carries a warning that DeepSeek plans to raise overall API pricing in the near future, with a significant increase expected, so treat today's figures as current rather than settled.

Are the DeepSeek-V4-Pro-0813 weights open?

Not as of 12 August. There is no DeepSeek-V4-Pro-0813 repository on Hugging Face, and the most recent upload on the deepseek-ai account is still DeepSeek-V4-Flash-0731 from 1 August. The April V4-Pro preview weights are up under the MIT licence, so if you self-host today you are running the preview, not the build the API now serves.

Should I use V4 Pro or V4 Flash?

Flash for almost everything, Pro when the task is genuinely hard. On DeepSeek's own Terminal Bench 2.1 numbers the Pro 0813 build scores 87.9 against 82.7 for Flash 0731, so the flagship is ahead again after two weeks of being behind. That is 5.2 points for 3.1x the token price and a fifth of the concurrency, since Pro is capped at 500 concurrent requests against 2500 for Flash.

Tags: aiapideepseekdev-toolsllmnewsopen-weights
Share197Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.