• Latest
  • Trending
  • All
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

7 September 2026
Answer card stating that Mullvad announced on 3 September 2026 that it is shutting down its public encrypted domain name system servers on 2 November 2026 and sponsoring the Quad9 Foundation instead, with 194.242.2.2 and its five sibling addresses all going away, and virtual private network customers unaffected.

Mullvad’s DNS servers go dark on 2 November, and Quad9 blocks no ads

5 September 2026
OpenAI announcement image for GPT-6 Astra, a spiral galaxy of white, blue and amber points of light curling around a bright core on a near black star field.

GPT-6 Astra lists at $10 and $50, 2.5x what GPT-5.6 Sol costs

6 September 2026
Google's official announcement image for the release, reading Introducing Gemini 3.8 Flash and 3.8 Flash Cyber in black type over a pale blue background with a blurred white chevron and the four colour Gemini spark below.

Gemini 3.8 Flash keeps the price and the 1 January cliff

3 September 2026
Answer card stating that Anthropic announced Enterprise Frontier Safeguards on 1 September 2026, that activity data used for misuse monitoring moves into cloud storage the customer controls under the customer own encryption keys, that Anthropic charges nothing for the feature while the cloud provider bills storage and egress, and that the phased rollout starts later in autumn 2026 with interim zero data retention on Fable 5 and Fable 5.1 for eligible customers.

Anthropic moves retention into your own cloud, for 30 days

3 September 2026
Official Google diagram of a client connection in three numbered steps: a DNS lookup with a query and an address, a TLS ClientHello and ServerHello, then a content exchange with a website. A callout on the DNS step reads 25% of global web traffic is now protected by encrypted DNS, and a callout beside an Android phone on the ClientHello step reads Android 17 supports ECH GREASE by default.

Android 17 hides the SNI, not your DNS or destination

3 September 2026
Still frame from the Claude Fable 5.1 launch video showing model-designed protein binders in orange docked against twelve grey target proteins, rendered as ESMFold2 structure predictions.

Claude Fable 5.1 breaks forced tool use, cuts cache 75%

1 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Tuesday, September 15, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

by stephane
12 September 2026
in Dev
0
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.
498
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

On Wednesday DeepSeek told its API users that V4 Pro had four days left. On Thursday it didn't. In between, on 10 September 2026, the company shipped DeepSeek-V4.1-Flash, a 552B mixture of experts with a new encoder-decoder layout, native image input and MIT weights, and renamed the endpoint to deepseek-flash. We've spent two days reading the pricing page and the change log against the model card, because they don't quite agree, and the disagreements are the useful part.

The short answer

DeepSeek-V4.1-Flash is live as deepseek-flash at $0.15 per million input tokens and $0.60 output off-peak, double that at peak, with cache hits at $0.003. The old deepseek-v4-flash and deepseek-v4-flash-vision-exp names are retired and temporarily route to it. On 10 September DeepSeek said every deepseek-v4-pro call would be routed to V4.1-Flash from 14 September at Flash rates. On 11 September the pricing page and change log gained a footnote saying V4 Pro stays on the API after 14 September with billing unchanged, in response to user demand. The launch post still carries the original sentence.

$0.15per 1M input off-peak, $0.60 out, cache hit $0.003
8B / 16Bactive parameters on input and output, of 552B
1 daybetween the V4 Pro retirement notice and its reversal
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.
Two DeepSeek pages, one day apart, saying different things about V4 Pro.

What shipped, and what your model name does now

The architecture is the actual news, and it's a real change rather than another re-post-training. V4.1-Flash is a 40 layer Transformer split into a 20 layer causal encoder and a 20 layer decoder. The decoder's global KV cache is projected from the encoder's final hidden states instead of being built layer by layer, which is how DeepSeek gets to 8B active parameters per token on prefill and 16B on decode, out of a 552B backbone. There's a 196B Engram conditional memory on top of that, plus a vision encoder trained from scratch, so the Hugging Face checkpoint reports 763B parameters in total and weighs about 510 GB across 88 files, most of it already in 8 bit. The card lists 384 routed experts per layer with 6 active, 45T training tokens, a 1M context, and a reasoning effort setting the card describes as an integer from 1 to 100, where the V4 API offered low, high and max. Weights are MIT, ungated, with vLLM and SGLang recipes.

On the API side, the name changed. Set model to deepseek-flash. The two previous names still work, but the models behind them are gone, and DeepSeek's own word for the routing is "temporarily". If you were on the vision experimental model from the Flash line we tracked through July, you're now on a different model with a different price and you didn't opt in, which has happened on this API more than once this year already. A minimal call looks like this:

Linux
curl https://api.deepseek.com/chat/completions -H "Authorization: Bearer $DEEPSEEK_API_KEY" -H "Content-Type: application/json" -d '{"model":"deepseek-flash","messages":[{"role":"user","content":"Which model are you?"}]}'

Vision is native on deepseek-flash and not supported on deepseek-v4-pro, per the pricing page's feature matrix. Both keep the 1M context and the 384K output ceiling. Concurrency caps are unchanged: 2,500 for Flash, 500 for Pro.

A price cut, unless you remember July

DeepSeek's launch post says the new architecture lets it serve more users for less and that it's passing the savings on. That's true against last week. Off-peak, V4.1-Flash costs $0.003 per million tokens on a cache hit, $0.15 on a miss and $0.60 on output. V4 Flash, since the peak and off-peak split arrived on 16 August, cost $0.007, $0.22 and $0.66 in the same slots. So input dropped by about a third and output by 9%. Peak hours, which are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, cost exactly double: $0.006, $0.30 and $1.20.

Horizontal bar chart of DeepSeek API output prices in dollars per million tokens at off peak rates, showing V4 Pro at 1.98 dollars, V4 Flash between 16 August and 9 September 2026 at 0.66 dollars, V4.1 Flash from 10 September at 0.60 dollars, and V4 Flash before 16 August at 0.28 dollars, with a note that peak hour rates are exactly double and the pre August price had no peak window.
Cheaper than the August rate. Not close to the July one.

Now go back one more step. Until 16 August, V4 Flash was a flat $0.0028, $0.14 and $0.28, no peak window at all. Against that, the shiny new model is dearer on every line, and more than double on output even at the discounted hours. Put it on the workload we used in July, an agent loop pushing 10 million input and 2 million output tokens a day with 90% of the input hitting cache. On the July rates that was about $0.73 a day. On V4.1-Flash off-peak it's $0.15 for the misses, roughly $0.03 for the hits and $1.20 for output, so $1.38. Run it at peak and you're at $2.75. That's our arithmetic on published rates, not a DeepSeek figure, and it's the number I'd want in front of anyone who reads "lower API prices" in the announcement.

The part that may actually save you money isn't on the price list. The global KV cache is down to 890 bytes per token, about a quarter of V4 Flash, and the persistent cache footprint on disk is about an eighth. DeepSeek says cache hit charges are often a large share of agent costs and that compressing the cache is what let it cut them. If you self-host, that's a straight HBM saving. If you use the API, it's the reason the cache hit price is $0.003 rather than $0.007, and it only pays off if your prefix is stable enough to hit. Check your hit rate before you check anything else.

Pro was retired on Wednesday and un-retired on Thursday

The launch post is blunt. Tests by multiple parties, it says, put V4.1-Flash ahead of V4 Pro on performance, cost, speed and total runtime, so "we're phasing out V4-Pro", and from 04:00 UTC on 14 September every deepseek-v4-pro request would be routed to V4.1-Flash at Flash rates until a V4.1-Pro launches. Trade press ran it as the flagship being killed. Then, on 11 September, footnote two on the pricing page and a new paragraph in the change log said that in response to user demand DeepSeek will keep serving V4 Pro after 14 September with billing unchanged, and will give notice of any future change. As we write this, the news page still shows the routing sentence. Nobody has edited it.

Why the demand? Look at the rows where Flash doesn't win. On DeepSeek's own table V4.1-Flash beats V4 Pro 0813 across the agentic block: 90.6 against 87.9 on Terminal-Bench 2.1, 74.2 against 62.7 on DeepSWE v1.1, 31.2 against 12.4 on Terminal-Bench 4.0. But Pro still leads on GPQA Diamond, 92.4 to 90.9, and on the text-only slice of HLE, 42.7 to 39.1. A 16B-active decoder is a brilliant coding agent and a slightly worse encyclopaedia, which is roughly what you'd predict. Someone with a knowledge-heavy workload asked to keep paying $1.98 a million for the bigger brain, and DeepSeek said fine.

Horizontal bar chart of Terminal-Bench 4.0 scores from the DeepSeek V4.1 Flash model card, showing Claude Opus 5 at 51.8 percent, GPT-5.6 Sol at 39.9 percent, GLM-5.3 at 37.9 percent, DeepSeek V4.1 Flash at 31.2 percent, DeepSeek V4 Pro at 12.4 percent and DeepSeek V4 Flash at 7.0 percent.
DeepSeek's own table. A 2.5x jump over Pro, and 20 points short of the top.

Two cautions on those scores, both printed by DeepSeek itself. Everything was run at the maximum effort setting, and the code agent rows use DeepSeek Harness in minimal mode, which the card says produced 90.6 on Terminal-Bench 2.1 where Claude Code got 88.0 and Codex 84.1 with the same model. Your harness picks your number. And the table that puts Flash at 74.2 on DeepSWE, a hair over Claude Opus 5 at 74.0, also puts it at 31.2 on Terminal-Bench 4.0 against 51.8. Both rows are honest. Only one of them will be quoted.

What we'd do this week: if you're on Pro for coding agents, try deepseek-flash on the calls that hurt, because the table says it's better there at a third of the output price. If you're on Pro for anything that leans on world knowledge, stay, and put a calendar note on the wording "we will provide further notice". Honestly, I read the one-day reversal as a good sign, a vendor that listens, but I might be wrong about how long it holds. A company that planned to route its flagship into a cheaper model on four days' notice can plan it again, and the notice period is the thing to watch, not the model.

Sources

DeepSeek, DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient, 10 September 2026 (architecture summary, the deepseek-flash name, the KV cache ratios, the 14 September routing statement, the pricing effective date). DeepSeek, Models & Pricing, read 12 September 2026 (the full peak and off-peak table, the feature matrix, concurrency limits, and footnote two on V4 Pro continuing after 14 September). DeepSeek, API change log, entries of 13 August, 21 August and 10 September 2026 (the 16 August peak and off-peak switch, the vision experimental model, the reversal paragraph). DeepSeek, DeepSeek-V4.1-Flash model card on Hugging Face, 10 September 2026 (the benchmark table, the harness comparison, the MIT licence, the parameter counts, and the 510 GB repository size from its file listing). PANews, DeepSeek V4 Pro API Call Service Will Not Be Discontinued, 11 September 2026 (independent timing of the reversal). VentureBeat, DeepSeek-V4.1-Flash debuts with $0.003/1M off-peak cached-input rate, 10 September 2026 (independent confirmation of the pricing and the original Pro routing plan). The pre-16 August prices are from our own July and August coverage of the pricing page. The daily cost figures are our arithmetic on the published rates.

Frequently asked questions

What is DeepSeek-V4.1-Flash?

DeepSeek's first model on a new architecture, released on 10 September 2026. It's a 552B parameter mixture of experts built as a 20 layer causal encoder followed by a 20 layer decoder, which lets it activate 8B parameters per token on input and 16B on output. It takes images natively, has a 1M token context, and the weights are on Hugging Face under the MIT licence. DeepSeek calls it the smallest model in the new family, which implies larger ones are coming.

How much does DeepSeek V4.1-Flash cost?

Per million tokens, off-peak: $0.003 on a cache hit, $0.15 on a cache miss and $0.60 for output. Peak hours, 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, cost exactly double. That's lower than the V4 Flash rate in force since 16 August ($0.22 in, $0.66 out off-peak) and higher than the flat $0.14 and $0.28 V4 Flash cost before that date.

Is DeepSeek V4 Pro being discontinued?

Not any more. The 10 September launch post said all deepseek-v4-pro requests would be routed to V4.1-Flash at Flash rates from 04:00 UTC on 14 September until a V4.1-Pro launched. On 11 September DeepSeek added a note to its pricing page and change log saying it will continue to serve V4 Pro after 14 September with billing unchanged, in response to user demand, and will give further notice of any change. The launch post hasn't been updated to match.

Do I need to change my code for deepseek-flash?

Eventually, yes. The new model name is deepseek-flash. The old names deepseek-v4-flash and deepseek-v4-flash-vision-exp still resolve, but DeepSeek describes that routing as temporary and the models behind them are retired, so anything still using them is already talking to V4.1-Flash at V4.1-Flash prices. Change the string now and re-run your evals, because the model behind it changed on 10 September whether you did or not.

Can I run DeepSeek-V4.1-Flash locally?

The weights are MIT licensed and ungated, and the model card ships vLLM and SGLang commands. The repository is about 510 GB across 88 files, mostly stored in 8 bit, and the card recommends a 1M context with at least 256K output tokens, so this is a multi-GPU server job rather than a workstation one. The 890 byte per token KV cache is the part that makes long contexts tractable at all.

Tags: API pricingdeepseekllmnewsopen-weights
Share199Tweet125
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.