• Latest
  • Trending
  • All
Official Qwen architecture diagram for Qwen3.8-Flash-Next, showing input tokens feeding a vocabulary embedding and a separate n-gram embedding layer at layer 2, a stack of Gated DeltaNet layers punctuated by Qwen Sparse Attention layers in a three to one hybrid block, Gated Residual read and write gates around each MoE block, and MTP modules beside the prediction head.

Qwen3.8-Flash-Next runs 6B active and weighs 173 GiB

3 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Sunday, September 20, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Qwen3.8-Flash-Next runs 6B active and weighs 173 GiB

by stephane
3 September 2026
in Dev
0
Official Qwen architecture diagram for Qwen3.8-Flash-Next, showing input tokens feeding a vocabulary embedding and a separate n-gram embedding layer at layer 2, a stack of Gated DeltaNet layers punctuated by Qwen Sparse Attention layers in a three to one hybrid block, Gated Residual read and write gates around each MoE block, and MTP modules beside the prediction head.
504
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Somebody on the team read "6B activated" and asked which of our boxes could run it. Fair question, wrong number: the FP8 checkpoint for Qwen3.8-Flash-Next is about 173 GiB of safetensors, and the BF16 one is 335 GiB. Six billion is what fires per token, not what you have to keep somewhere. Alibaba put the model out on 26 August 2026 as an open-weight preview of the architecture behind Qwen4, and honestly the engineering in it is the most interesting thing we have read all month. Two details keep getting flattened in the coverage, though. The licence is not Apache 2.0, whatever a few launch write-ups said. And the model you call on the API is named Qwen3.8-Flash, which is not the same artefact as the one on Hugging Face.

The short answer

Qwen3.8-Flash-Next is a preview of the Qwen4 architecture, open-weight since 26 August. Genuinely clever design, real benchmark wins over a model three times its size. Just don’t confuse it with the Qwen3.8-Flash you call on the API, and read the licence before you build a product on it.

125Btotal, with 6B activated per token
173 GiBthe FP8 checkpoint on Hugging Face
NotApache 2.0, whatever you read
Answer card: Qwen3.8-Flash-Next released 26 August 2026 with 125 billion parameters, 6 billion activated per token, a 51 billion parameter n-gram embedding table and a 4 billion parameter MTP head, licensed under Qwen Community 1.0 rather than Apache 2.0, with the production API version named Qwen3.8-Flash priced at 0.16 dollars per million input tokens and 0.47 dollars per million output tokens, and a context ceiling of one million tokens with YaRN from 262,144 native.
The four numbers that matter, and the two names.

Two names, one release

This trips people up within about thirty seconds of the announcement, so let’s clear it first.

The weights on Hugging Face are Qwen/Qwen3.8-Flash-Next, with an FP8 variant beside them. The thing you hit on Qwen Cloud is Qwen3.8-Flash. Qwen’s own model card puts it plainly:

Qwen3.8-Flash is the official version based on Qwen3.8-Flash-Next with more production features, e.g., 1M context length by default, official built-in tools.

So the API model is a superset, not a mirror. If you benchmark the open weights locally and then ship against the hosted endpoint, you are not testing the same system. That matters more than usual here, because the context handling is one of the things that differs: the open release is 262,144 tokens natively and stretches to a million with YaRN, while the hosted one starts at a million.

Pricing on Qwen Cloud is $0.16 per million input tokens and $0.47 per million output. OpenRouter’s listing reads $0.15 in, same $0.47 out. Nobody has explained the cent.

The architecture is the actual story

Official Qwen architecture diagram for Qwen3.8-Flash-Next, showing input tokens feeding a vocabulary embedding and a separate n-gram embedding layer used at layer 2 only, a stack of Gated DeltaNet layers punctuated by Qwen Sparse Attention layers in a repeating three to one hybrid block, Gated Residual read and write gates wrapping each MoE block, and MTP modules beside the prediction head.

Image: Qwen (Alibaba Group), from the Qwen3.8-Flash-Next model card

Forty-eight layers, arranged as twelve repetitions of three Gated DeltaNet blocks followed by one Qwen Sparse Attention block. Every one of those blocks feeds a mixture of experts with 512 experts, of which ten routed plus one shared fire per token. Expert intermediate dimension is a slim 640. Hidden dimension across the model is 2560, which is small for something with 125B parameters in it, and that’s the point: the width lives in the expert count, not the tensor shapes.

Then there’s the n-gram table. Twenty million bigram and trigram entries, 51B parameters, consulted at layer 2 only. Qwen’s framing is that embeddings scale parameters without scaling compute and are easier to offload than MoE weights are. Which reads as a polite way of saying they found somewhere cheaper than HBM to keep half the model.

I might be wrong about how well that holds up under a real serving load. The idea is sound, the memory hierarchy is not always as forgiving as an architecture diagram makes it look.

The 6B figure, read properly

Bar chart comparing one H100 at 80 GB of HBM and one H200 at 141 GB against the Qwen3.8-Flash-Next FP8 checkpoint at about 173 GiB and the BF16 checkpoint at about 335 GiB, measured as the sum of the Hugging Face repository files on 30 August 2026.
We summed the repo files ourselves. Neither number fits on one card.

Six billion parameters activate per token. That’s a compute claim, and it’s a good one. It is not a memory claim, and the gap between the two is where the disappointment usually lands.

We pulled the file sizes off the Hugging Face API on 30 August: 144 files summing to 172.8 GiB for the FP8 repo, 335.3 GiB for BF16. Add the KV cache for whatever context you actually intend to use, add framework overhead, and this is a small cluster or a very well equipped single node. Not a workstation.

Compare that to GLM-5.3-Flash, which Z.ai shipped the same day at 320B total and 18B active. Different tradeoff, same conclusion for anyone hoping to self-host on a desk.

What the benchmarks say, losses included

Credit where it’s due: Qwen published a table that includes the places it comes second.

It beats Qwen3.7-Plus, a 397B model with 17B active, almost everywhere. SWE-bench Pro 62.5 against 55.8. CoWorkBench 73.9 against 65.1. JobBench 55.7 against 27.6, which is a gap wide enough to make us suspicious of the harness rather than impressed by the model. Same instinct on DeepSWE 1.1, where Flash-Next posts 58.7 and Qwen3.7-Plus posts 16.5. Numbers that far apart usually mean the older model was failing to complete the task format, not that it’s fifteen times worse at software engineering.

The losses are more informative. On repo-level generation, NL2Repo-Bench, DeepSeek-V4-Flash-0731 takes it 54.2 to 48.1. On HLE, Claude Opus 4.6 Max wins 40.0 to 35.9. And on Agents’ Last Exam the Pass@1 goes to DeepSeek, 25.2 to 24.3.

One thing the table doesn’t do is compare against anything newer than Opus 4.6. Opus 4.8 and Opus 5 both shipped before 26 August. Their absence is a choice.

The licence is the part to actually read

Checklist of the Qwen Community License 1.0 terms: commercial use, modification and distribution are permitted and internal deployment is fine, running a Model as a Service or AI Work Assistant business requires a separate licence from Qwen first, MaaS covers giving third parties inference or fine-tuning access via API, relaying requests to a third-party hosted copy is excluded, products above 100 million monthly active users or 20 million dollars monthly revenue must display the model name, and the licence is not Apache 2.0.
Open weights and open serving are not the same permission.

The repo metadata says license: other, license_name: qwen-community-1.0. Not Apache 2.0. Several outlets reported Apache 2.0 on launch day and it isn’t a small distinction.

You can use it commercially, modify it, redistribute it, run it inside your own company. What you cannot do without asking Qwen first is run a Model as a Service or an AI Work Assistant business on it. The licence defines MaaS as giving third parties access to inference or fine-tuning via API, and explicitly carves out relaying requests to a third-party hosted copy. So reselling an endpoint you host yourself needs a separate licence. Pointing your users at somebody else’s hosted Qwen doesn’t.

Above 100,000,000 monthly active users or US$20,000,000 monthly revenue, you also have to display the model name prominently in the product.

So when would we reach for it

If you’re paying for a hosted frontier model to do long agentic runs, $0.16 per million in is worth an afternoon of evaluation, particularly on the office-work and tool-use side where CoWorkBench and Toolathlon flatter it. If you want to self-host, budget for a node and not a card, and take the FP8 checkpoint.

And if you’re building a product that sells inference to other people, read the licence before you read the benchmarks. That order saves time.

Sources

Model card and licence: Qwen/Qwen3.8-Flash-Next on Hugging Face and the Qwen Community License 1.0. Official announcement: Qwen3.8-Flash-Next, A New Architecture, Towards Ultimate Cost-Efficiency and the QwenLM repository. API listing: qwen/qwen3.8-flash on OpenRouter. Coverage: TechNode and MarkTechPost. Checkpoint sizes are our own sum of the repository files via the Hugging Face API on 30 August 2026.

Frequently asked questions

What is the difference between Qwen3.8-Flash-Next and Qwen3.8-Flash?

Flash-Next is the open-weight release on Hugging Face and ModelScope, published on 26 August 2026 under the Qwen Community License 1.0. Qwen3.8-Flash is the managed version on Qwen Cloud, described by Qwen as based on Flash-Next but with production features on top, including a 1M context length by default and official built-in tools. Same family, different artefacts, and the API one is the one with a price attached.

How much does Qwen3.8-Flash cost?

Qwen lists $0.16 per million input tokens and $0.47 per million output tokens on Qwen Cloud. The OpenRouter listing for qwen/qwen3.8-flash reads $0.15 in and $0.47 out on 30 August 2026, with a 1,000,000 token context and a 131,072 token completion ceiling. The one cent gap between the two input rates is not something either side has explained.

Is Qwen3.8-Flash-Next Apache 2.0?

No. Several launch write-ups said Apache 2.0 and they were wrong. The Hugging Face repo metadata reads license: other with license_name: qwen-community-1.0, and the licence file is the Qwen Community License 1.0. It permits commercial use and redistribution, but it requires a separate licence from Qwen if you run a Model as a Service or an AI Work Assistant business.

Can I self-host it on one GPU?

Not on one card. The FP8 repo on Hugging Face sums to roughly 173 GiB and the BF16 repo to about 335 GiB, before any KV cache. The 6B activated figure describes compute per token, not residency. Qwen does note that the n-gram embedding table offloads more gracefully than mixture of experts weights do, so some of the 51B can live off the accelerator, and it points production users at SGLang or vLLM rather than plain Transformers.

Is it better than DeepSeek V4-Flash or Claude Opus?

Depends what you measure, and Qwen published its own losses. On the vendor table Flash-Next takes SWE-bench Pro at 62.5 against 56.0 for DeepSeek-V4-Flash-0731 and 53.4 for Claude Opus 4.6 Max. It loses NL2Repo-Bench, 48.1 against 54.2 for DeepSeek, and it loses HLE, 35.9 against 40.0 for Opus 4.6 Max. Worth noting the comparison set stops at Opus 4.6, so 4.8 and Opus 5 are absent.

Tags: aillmnewsopen-weightsqwenself-hosting
Share202Tweet126
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.