• Latest
  • Trending
  • All
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

7 September 2026
Answer card stating that Mullvad announced on 3 September 2026 that it is shutting down its public encrypted domain name system servers on 2 November 2026 and sponsoring the Quad9 Foundation instead, with 194.242.2.2 and its five sibling addresses all going away, and virtual private network customers unaffected.

Mullvad’s DNS servers go dark on 2 November, and Quad9 blocks no ads

5 September 2026
OpenAI announcement image for GPT-6 Astra, a spiral galaxy of white, blue and amber points of light curling around a bright core on a near black star field.

GPT-6 Astra lists at $10 and $50, 2.5x what GPT-5.6 Sol costs

6 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Friday, September 18, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

by stephane
18 September 2026
in Dev
0
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.
492
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

We've got a folder of meeting recordings nobody transcribes because the audio tariff on every omni model made it a budget line. Late on 17 September 2026 the Qwen team published Qwen3.8-Omni-Flash, and the Model Studio pricing page now lists it at $0.15 per million input tokens whether the tokens came from text, an image, a video or an hour of audio. That's the same rate as the text-only qwen3.8-flash. So we read the launch post, the pricing table and the API docs, then went looking for the parts of the announcement you can actually download. Some of them you can't.

The short answer

Qwen3.8-Omni-Flash is an API model on Alibaba Cloud Model Studio. It accepts text, images, audio and video in a 1M-token context and returns text only. International regions bill $0.15 per million input tokens for every modality, $0.016 on a cache hit and $0.47 per million output. The previous Omni-Plus charged $11 for audio input, so the cut is 98.6% on that line. There are no open weights, the Qwen-Live Harness repository the post links to returns 404, and the Realtime variant has no row on the pricing page yet. The benchmark table is the lab's own, and it loses three of the four agentic rows to Gemini 3.8 Flash.

$0.15 / $0.47per million tokens in and out, all modalities, international regions
1M tokenscontext, with an hour of audio-visual input handled natively per the launch post
Sept 2025the last time the team published Omni weights
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.
The price is the story. The weights, the harness and the speech output are the footnotes you'll hit first.

What $0.15 a million actually buys

Start with the table, because the launch post talks in percentages and the pricing page talks in dollars. In the international regions (Singapore, Hong Kong, Tokyo, Frankfurt, Virginia) qwen3.8-omni-flash has one input column: USD 0.15 per million tokens. No separate audio rate, no separate video rate. Cache hits are $0.016 and text output is $0.47. Compare that with qwen3.5-omni-plus on the same page, still listed at $1.4 for text, image and video input, $11 for audio input and $8.3 for text output. The post claims an audio-input cut "of more than 98%" per hour; $11 to $0.15 per million tokens is 98.6%, so the claim checks out against the lab's own tariff. The older qwen3.5-omni-flash sat at $3 for audio. Even that is twenty times what the new model charges.

Two things the price doesn't include. Speech output, first. The docs are explicit that this model returns text, and that generated speech stays on Qwen3.5-Omni or on qwen3.8-omni-flash-realtime, the WebSocket and WebRTC variant. The Realtime model is in the launch post with latency numbers (about 591 ms to first token and 978 ms to the first audio packet on a six-second clip) but when we read the pricing page on 18 September there was no row for it. Second, the free quota. Model Studio's note says the Omni free tiers exist only in Singapore, and the 3.8 row we saw had no quota column at all. Budget for pay-as-you-go from the first call.

What you do get is a 1M-token window (the post says an hour of audio-visual input is handled natively, and the docs cap a Base64 upload at 10 MB), 113 recognised languages and dialects, stereo and four-channel spatial audio, function calling, and an OpenAI-compatible surface on both Chat Completions and Responses. The only built-in Responses tool is web search. Reasoning is on by default at reasoning_effort="xhigh", with medium, low and none below it, and preserve_thinking is enabled out of the box, which means prior reasoning rides along in the next turn and gets billed as input. Turn it down before you point it at a two-hour recording, or the context fills faster than you'd expect.

curl
curl https://dashscope-intl.aliyuncs.com/compatible-mode/v1/chat/completions -H "Authorization: Bearer $DASHSCOPE_API_KEY" -H "Content-Type: application/json" -d '{"model":"qwen3.8-omni-flash","reasoning_effort":"low","messages":[{"role":"user","content":[{"type":"input_audio","input_audio":{"data":"https://example.com/meeting.wav","format":"wav"}},{"type":"text","text":"List the action items with who owns each."}]}]}'
Horizontal bar chart of audio input prices in dollars per million tokens on Alibaba Cloud Model Studio international regions, showing Qwen3.5-Omni-Plus at 11 dollars, Qwen3-Omni-Flash at 3.81 dollars, Qwen3.5-Omni-Flash at 3 dollars and Qwen3.8-Omni-Flash at 0.15 dollars, with a note that the new model bills text, image, video and audio input at the same rate and returns text only.
Four Omni models on one pricing page. The new one bills audio like text, which is the part the launch post is right to shout about.

The lab's table, read row by row

The headline claim is an average gain of more than 25% over Qwen3.5-Omni-Plus across 29 evaluations, and audio-visual performance "close to" Gemini 3.8 Flash. Read the rows and it's more uneven than that, in both directions. On the four agentic rows, the new model wins WildClawBench-MM at 71.0 against 58.9 and edges UniClawBench 69.6 to 69.0, then loses AgenticVBench 36.8 to 45.0 and OmniGAIA 74.0 to 78.6. On plain speech recognition it's behind its own predecessor: FLEURS multilingual WER is 9.3 against 7.2 for Omni-Plus and 7.9 for the Gemini model, and VoiceBench slips from 92.9 to 91.6. Honestly, for a model marketed on audio, a worse WER than the model it replaces is the number we'd want explained.

The row that is genuinely new is multi-speaker meetings. On AliMeeting the diarisation error rate goes from 88.1 to 3.4 and the concatenated WER from 89.6 to 17.2. Those old numbers are so bad they say the previous model simply couldn't separate speakers, so this is a capability arriving rather than improving. Same pattern on AISHELL-4 and MagicData-RAMC. If your use case is "who said what in an hour-long call", that's the row to care about, and it's the one where the Gemini and Muse Spark 1.2 columns look worst too.

One more number we liked. In agent mode (the post uses its own Qwen Code harness) the model scores 67.8 on OmniVideoBench instead of 63.4, while spending 79,117 tokens per query instead of 145,736. Watching less of the video and answering better is the interesting result, more than any single score. The text rows, including a 63.3 on SWE-bench Pro, were run on a version of the benchmark the lab says it corrected and re-ran itself, so we wouldn't line those up against anyone else's published score. Nobody outside the lab had rerun any of this when we checked on 18 September, which is normal for day one and worth remembering all the same.

What you can't download yet

The launch post links two open-source companions. Qwen-MM-Plugins is real: a public GitHub repository under Apache 2.0, pushed again on the morning of the launch, with a video-to-notes plugin, a film commentary one, and a skill creator that turns a screen recording into an agent skill file, among others. The other one, Qwen-Live Harness, is described as "open-sourced" with an npm install -g qwen-live-harness one-liner. The package exists on the npm registry (0.4.2, published 15 September) and its repository field points at github.com/QwenLM/Qwen-Live-Harness. That URL returned 404 every time we tried it on 18 September, and people in the Hacker News thread hit the same wall. I might be wrong about why; a private repo waiting on a flip is the boring explanation. But as of today the harness is a package you can install without any source you can read, which isn't what open-sourced means.

Then the weights. Nothing under the Qwen organisation on Hugging Face for 3.8-Omni. The last Omni checkpoints the team published were the Qwen3-Omni-30B-A3B family in September 2025, a full year ago. Since then the omni line has been API-only, while the text side kept shipping: Qwen3.8-Flash-Next landed with weights three weeks ago and the 27B is Apache 2.0. Our read is that omni is where the team has decided to keep the moat. The post doesn't mention weights at all, which for a lab that usually says "coming" when they're coming is its own kind of answer.

So who is this for? Anyone whose audio bill was the reason they didn't build the thing. At $0.15 a million the audio line stops being the reason not to build it, and you're not paying a modality premium for the video track either. Against Gemini 3.8 Flash it isn't a clean win on the lab's own table, and if you need the model to talk back, the honest comparison is with the Gemini 3.8 Live models on the one side and an unpriced Realtime endpoint on the other. We'd run the meeting folder through it this week. We wouldn't plan a product on the harness until the repository loads.

Checklist of what the Qwen3.8-Omni-Flash launch does and does not ship on 18 September 2026, listing that the API is live in six Model Studio regions with OpenAI compatible Chat Completions and Responses, that Qwen-MM-Plugins is public on GitHub under Apache 2.0, that the model returns text only with no speech output, that the Qwen-Live Harness repository linked from the blog returns 404 although the npm package exists, and that no Qwen3.8-Omni weights exist on Hugging Face.
What we could reach from the launch page on 18 September. Two green, three red, and the red ones are the ones people will ask about.

Sources

Qwen team, Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery., September 2026 (the 1M context, the 29-evaluation average, the benchmark tables, the OmniVideoBench token counts, the Realtime latency figures, the reasoning_effort levels, the harness and plugin links). Alibaba Cloud, Model Studio model inference pricing, read 18 September 2026 (the $0.15, $0.016 and $0.47 rows, the Qwen3.5-Omni-Plus and Omni-Flash rows, the free quota note). Alibaba Cloud, Qwen-Omni documentation, read 18 September 2026 (text-only output, the six regions, the 10 MB Base64 limit, the spatial audio support, the 113 languages, the Responses tool scope). GitHub, QwenLM/Qwen-MM-Plugins (Apache 2.0, last push 18 September 2026) and the 404 at QwenLM/Qwen-Live-Harness, both checked 18 September 2026. npm registry, qwen-live-harness, versions 0.4.1 and 0.4.2 published 14 and 15 September 2026. Hugging Face, the Qwen organisation, checked 18 September 2026 (no Qwen3.8-Omni repository; Qwen3-Omni-30B-A3B dated September 2025). Hacker News, Qwen 3.8 Omni Flash, 17 September 2026 (the harness 404 reports).

Frequently asked questions

How much does Qwen3.8-Omni-Flash cost?

In Alibaba Cloud Model Studio's international regions it's USD 0.15 per million input tokens for text, image, video and audio alike, USD 0.016 per million on a cache hit and USD 0.47 per million output tokens. The Beijing region is priced separately. The Realtime variant had no price on the page when we read it on 18 September 2026.

Are Qwen3.8-Omni-Flash weights available on Hugging Face?

No. As of 18 September 2026 there's no Qwen3.8-Omni repository under the Qwen organisation, and the launch post doesn't mention open weights. The most recent Omni checkpoints the team published are the Qwen3-Omni-30B-A3B models from September 2025.

Can Qwen3.8-Omni-Flash generate speech?

Not this model. The docs list its output as text only. For spoken responses you use Qwen3.5-Omni over HTTP or qwen3.8-omni-flash-realtime over WebSocket or WebRTC, which the launch post describes but which isn't yet on the pricing page.

Is Qwen3.8-Omni-Flash better than Gemini 3.8 Flash?

On the lab's own table it wins some rows and loses others. It leads WildClawBench-MM 71.0 to 58.9 and most of the multi-speaker meeting rows, and trails on AgenticVBench, OmniGAIA and multilingual speech recognition. Nobody independent had rerun the numbers on launch day.

Does the Qwen-Live Harness work?

The npm package qwen-live-harness exists at version 0.4.2, but the GitHub repository it points to returned 404 on 18 September 2026, so there's no readable source yet. Qwen-MM-Plugins, the other companion project, is public under Apache 2.0.

Tags: Alibaba CloudapiaudioLLM pricingnewsomnimodalqwenQwen3.8
Share197Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.