• Latest
  • Trending
  • All
Answer card: Alibaba published the Qwen3.8-27B weights on 14 August 2026 under Apache 2.0 as a dense native vision language model with 262,144 tokens of native context, while the 2.4 trillion parameter flagship weights published two days earlier carry a custom qwen3.8-max licence and are text only.

Qwen3.8-27B ships Apache 2.0 with vision, the 2.4T doesn’t

15 August 2026
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Wednesday, September 23, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Qwen3.8-27B ships Apache 2.0 with vision, the 2.4T doesn’t

by stephane
15 August 2026
in Dev
0
Answer card: Alibaba published the Qwen3.8-27B weights on 14 August 2026 under Apache 2.0 as a dense native vision language model with 262,144 tokens of native context, while the 2.4 trillion parameter flagship weights published two days earlier carry a custom qwen3.8-max licence and are text only.
495
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Open the two model cards side by side and the surprise is which one you would rather self host. Qwen3.8-27B went up on Hugging Face on 14 August under Apache 2.0: dense, with a vision encoder nobody had been promised, 262,144 tokens of native context, and a reasoning effort dial you can actually turn down. The 2.4 trillion parameter flagship whose weights landed two days earlier carries a bespoke licence called qwen3.8-max, and its own card says multimodal inputs are not supported and thinking cannot be disabled. So the small one is the permissive one. It is also the multimodal one. We went looking for the catch, and the catch is mostly VRAM.

The short answer

Alibaba shipped two sets of Qwen3.8 weights inside three days. The 2.4T flagship came first, text only, under a bespoke licence. Then the 27B landed under Apache 2.0 with a native vision encoder. If you self host, the second one is the interesting release, and it is the one that fits on hardware you might already own.

Apache 2.0on the 27B, custom licence on the 2.4T
visionon the 27B only
~28 GBFP8 weights, one 48 GB card
Answer card: Alibaba published the Qwen3.8-27B weights on Hugging Face on 14 August 2026 under Apache 2.0, a dense native vision language model with 262,144 tokens of native context extensible to one million, while the 2.4 trillion parameter flagship weights published two days earlier carry a custom qwen3.8-max licence and are text only.
Two weight drops, two days apart, and the small one has the better terms.

What actually landed

Back on 4 August we wrote that Qwen3.8-Max was not open source, that it was API only, and that the weights were promised for the following week with no licence named. Both halves of that promise have now been kept, sort of.

The flagship weights went up around 12 August as Qwen/Qwen3.8-2.4T-A95B. Then on 14 August the 27B appeared, and it is the one that got the attention.

Here is the shape of it. Qwen3.8-27B is dense, not sparse. Sixty four layers, built as sixteen repeats of a block that mixes Gated DeltaNet with Gated Attention. Native context of 262,144 tokens, extensible to a million. And a vision encoder, which is the part nobody had been trailing: the card calls it a native vision language model that understands images and videos.

Side by side comparison of the two Qwen3.8 open weight releases: the 27B under Apache 2.0 with a native vision encoder and configurable reasoning effort, against the 2.4T mixture of experts flagship under the custom qwen3.8-max licence, text only and with thinking forced on.
Same family, same context window, and a licence field that diverges hard.

The licence is the line I would read first. Apache 2.0 on the 27B, which is boring in the best way: your legal team already has a position on it. The flagship card names qwen3.8-max instead, a bespoke licence, and reporting around that launch describes a revenue share aimed at large commercial users with neither the threshold nor the percentage published. Unpublished terms are not terms you can plan around, so I would treat the flagship weights as unusable for anything commercial until the actual text is out.

The bit that reads backwards

Normally the flagship gets the features and the small model gets the leftovers. Not here.

The 2.4T card is explicit: multimodal inputs are not supported, and thinking cannot be disabled. Both of those are real constraints. Forced thinking means every call pays the reasoning tokens whether the task needs them or not, which is exactly the problem we flagged on GLM-5.3 last week when Z.ai removed non-thinking mode. The 27B keeps the dial, with reasoning_effort at xhigh, medium or low.

So the downloadable flagship is a text only model that always thinks, and the downloadable 27B sees images and lets you turn the thinking down. Honestly, for most self hosted work that is the wrong way round from what the parameter counts suggest, and it is a good outcome.

The numbers Alibaba published

First real benchmark table for the Qwen3.8 family, incidentally. The July preview shipped with a tweet and no card at all.

On the 27B: SWE-bench Pro 61.7, Terminal Bench 2.1 73.0, LiveCodeBench v6 90.3, IFBench 79.5. On the vision side, OSWorld-Verified 84.3 and WebArena-Verified 64.8, with MathVision at 94.6.

The flagship posts higher where you would expect: SWE-bench Pro 67.7, Terminal Bench 2.1 86.6, GPQA Diamond 92.6, PaperBench 93.0. Six points of SWE-bench Pro between a 27B dense model and a 2.4 trillion parameter mixture of experts is a narrower gap than the parameter counts imply, and the usual caveat applies with full force. These are vendor numbers on vendor runs. Nobody outside Alibaba has reproduced any of them yet.

What I would take from the table is the shape rather than the digits. A 27B that lands in the low sixties on SWE-bench Pro and above eighty on OSWorld is a genuinely capable local agent model, and OSWorld is the computer-use benchmark, which is where the vision encoder earns its place.

Whether you can run it

This is the part that decides it for most people.

Bar chart of approximate VRAM for the Qwen3.8-27B weights at three precisions: about 56 GB at BF16 needing an 80 GB card, about 28 GB at FP8 fitting a 48 GB card, and about 14 to 17 GB at 4-bit fitting a 24 GB consumer card.
Weights only. The KV cache is the line item that will actually catch you out.

Alibaba publishes an FP8 build itself, Qwen3.8-27B-FP8, using fine grained quantisation at block size 128, and says the metrics come out nearly identical to the original. That build is compatible with Transformers, vLLM and SGLang. Community quantisations for llama.cpp, Ollama, LM Studio and Jan appeared within a day, as they always do.

Rough weight footprints: about 56 GB at BF16, about 28 GB at FP8, somewhere in the 14 to 17 GB range at 4-bit. Those are community estimates, not vendor figures, and they cover the weights only.

Then the cache. A 262k native window sounds like a gift until you price the KV cache that fills it, and at long context with a few concurrent requests the cache can outweigh the weights. If you are sizing a box on the 4-bit number because it fits a 24 GB card, budget for a short context or a single stream. That is the constraint people keep discovering after the hardware arrives, and it is the reason we keep saying the same thing about local models: the parameter count tells you almost nothing about what you need.

Would we run it

If you have a 48 GB card, yes, the FP8 build is the obvious thing to try this week. Apache 2.0, vision included, and a coding score that is respectable for the size.

If you were waiting on the 2.4T weights so you could self host the frontier model, the honest answer is that what shipped is not what the hosted product is. Text only, thinking forced, licence unpublished in its details. It exists, you can download it, and I am not sure who it is for yet.

And if you are picking a local model for agent work specifically, the OSWorld number plus the vision encoder is the argument for this one over a text only competitor of similar size. Worth a weekend. We will run it against Qwen3.7 on local hardware once there is something to compare beyond Alibaba’s own table.

Sources

Alibaba, Qwen3.8-27B-FP8 model card on Hugging Face, for the architecture, the Apache 2.0 licence, the 262,144 token native context, the native vision language description, the reasoning effort levels, the FP8 block size 128 quantisation and the benchmark table. Alibaba, Qwen3.8-2.4T-A95B model card, for the qwen3.8-max licence field, the 95B active parameters out of 512 experts, the statement that multimodal inputs are not supported and that thinking cannot be disabled, and its own benchmark figures. Yotta Labs, Qwen 3.8 27B specs and hardware requirements, for the VRAM estimates by precision. ExplainX, Qwen3.8-Max open weights are live, for the 12 August flagship weights date and the reporting on the revenue share clause, whose threshold and percentage are not published.

Frequently asked questions

What licence is Qwen3.8-27B under?

Apache 2.0. The licence field on the Qwen/Qwen3.8-27B-FP8 model card reads apache-2.0, which allows commercial use, modification and redistribution with attribution and no revenue share. That is a different licence from the flagship weights: the Qwen/Qwen3.8-2.4T-A95B card names a custom licence called qwen3.8-max instead.

Can Qwen3.8-27B read images?

Yes. Alibaba describes it on the card as a native vision language model that understands images and videos, and publishes vision results alongside the text ones, including 84.3 on OSWorld-Verified and 64.8 on WebArena-Verified. The 2.4T flagship weights cannot: that card states plainly that multimodal inputs are not supported.

How much VRAM does Qwen3.8-27B need?

Roughly 56 GB at BF16, about 28 GB at FP8 and somewhere around 14 to 17 GB at 4-bit, for the weights alone. The KV cache is extra and it grows with context length and concurrency, which matters here because the native window is 262,144 tokens. Alibaba does not publish VRAM figures itself, so treat these as the community estimates they are.

Is Qwen3.8-27B dense or mixture of experts?

Dense. The card describes 64 layers built as 16 repeats of a block mixing Gated DeltaNet and Gated Attention with feed forward layers. The 2.4T flagship is the sparse one, a fine grained mixture of experts with 95B parameters active per token out of 512 experts.

What context length does it support?

The card says 262,144 tokens natively, extensible up to 1,000,000. Both Qwen3.8 open weight releases quote the same pair of numbers, so context is not what separates them. The licence and the vision support are.

Tags: aialibaballmnewsopen-weightsqwen
Share198Tweet124
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.