• Latest
  • Trending
  • All
Answer card: Qwen 3.7 Max and GLM-5.2 are tied on coding benchmarks, but GLM is cheaper and open-weight while Qwen Max wins reasoning and long autonomous runs.

Qwen 3.7 Max vs GLM-5.2: benchmarks, price and the catch

22 June 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Monday, September 21, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Qwen 3.7 Max vs GLM-5.2: benchmarks, price and the catch

by stephane
22 June 2026
in Dev
0
Answer card: Qwen 3.7 Max and GLM-5.2 are tied on coding benchmarks, but GLM is cheaper and open-weight while Qwen Max wins reasoning and long autonomous runs.
497
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Alibaba dropped Qwen 3.7 Max in May 2026, and within days the coding crowd was lining it up against GLM-5.2, the open-weight model we already pulled apart here. Good instinct, because on raw coding they are basically the same model wearing different badges: 60.6 on SWE-bench Pro for Qwen against 62.1 for GLM, MCP-Atlas inside a single point, the kind of gap that vanishes on a re-run. So the benchmark race is a draw. The real decision is everything around it. GLM is cheaper and you can self-host it; Qwen 3.7 Max is API-only but pulls ahead on math, reasoning, and agents that stay coherent for hours. Here are the numbers side by side, the price gap, and the catch nobody puts in the headline.

The short answer

On coding, Qwen 3.7 Max and GLM-5.2 are a coin flip: under 1.5 points apart on every benchmark they share. They split everywhere else. GLM is cheaper and open-weight, so you can self-host it. Qwen 3.7 Max is API-only, but it wins math and reasoning and keeps an agent coherent for far longer.

<1.5 ptsgap on every shared coding test
$4.40 / $7.50output per 1M, GLM / Qwen
35 hQwen's autonomous run (GLM: ~8)
Answer card: Qwen 3.7 Max and GLM-5.2 tie on coding within 1.5 points; GLM is cheaper and open-weight, Qwen wins reasoning and 35-hour agent runs but is API-only.
The short version. The coding score is the least interesting number here.

What Qwen 3.7 Max actually is

Qwen 3.7 Max is Alibaba’s flagship, shipped on 21 May 2026, and it’s the closed one. Text-only, API-only, no weights to pull down (the open Qwen models are the smaller 3.6 line, like the dense Qwen3.6-27B). It carries a 1 million token context window, lands an Intelligence Index around 56.6 (fifth overall and the top-scoring Chinese model on that index), and posts the lowest hallucination rate of the current frontier at 22.9 percent.

Its real party trick is endurance. Qwen 3.7 Max can reportedly drive an autonomous agent for something like 35 hours before it loses the thread, against roughly 8 for GLM. Hold onto that number. It’s the single place these two models stop being interchangeable, and we’ll come back to it.

The coding race is a tie

Put them on the same coding benchmarks and you get a coin flip.

Grouped bar chart: GLM-5.2 vs Qwen 3.7 Max on SWE-bench Pro (62.1 vs 60.6), MCP-Atlas (77.0 vs 76.4) and HLE no tools (40.5 vs 41.4), all within about a point.
Three shared benchmarks, three near-ties. This is the whole coding story.

SWE-bench Pro: GLM 62.1, Qwen 60.6. MCP-Atlas: 77.0 to 76.4. HLE without tools, the one Qwen takes: 41.4 to 40.5. Every gap is under a point and a half, which is inside the noise. Run the suite again next week on a slightly different harness and the winner flips. The comparison that kicked this whole thing off called the two models “twins,” and honestly that’s the right word. If your job is closing GitHub issues and shipping pull requests, you cannot pick wrong here on quality, because there’s no quality gap to pick.

Which is exactly why the benchmark table is the least useful part of this comparison. Everyone fixates on it. It decides nothing.

Where they stop being twins

Two places, and they’re both real.

First, reasoning and math, where Qwen pulls clear air. GPQA Diamond at 92.4, a near-perfect 97.1 on the HMMT competition set, 91.6 on LiveCodeBench. If your work leans on hard, multi-step reasoning rather than patching an existing codebase, Qwen 3.7 Max is the sharper instrument, and it isn’t close. GLM has its own counter (it edges ahead with tools on HLE, 54.7, and on DeepSWE at 46.2), but pure reasoning is Qwen’s room.

Second, that endurance number. A 35-hour coherent autonomous run against GLM’s eight is not a rounding difference, it’s a different category of work. For an agent that has to grind through an overnight migration or a multi-day refactor without a human nudging it back on track, Qwen holds the plot together long after GLM has wandered off. If you’re building long-running agents, this one spec probably outweighs the entire benchmark table.

GLM’s answer is narrower but it lands for a lot of teams: it’s a touch ahead on the GitHub-style coding tests, it’s cheaper, and it’s open. That last word is the one that actually moves the decision.

The price, and the escape hatch

Bar chart of output price per million tokens: Qwen 3.7 Plus $1.60, GLM-5.2 $4.40, Qwen 3.7 Max $7.50, Claude Opus 4.8 $25, GPT-5.5 $30.
Same coding score, very different bill. And the cheapest two are the Chinese ones.

On output tokens, GLM runs $4.40 per million, Qwen 3.7 Max $7.50. So Qwen is about 1.7 times the bill for coding work that scores the same, which adds up at volume but won’t decide much on its own. The bigger split isn’t the number, it’s the door behind it. GLM-5.2 ships under an MIT license, so you can download the weights and run them on your own hardware, and at that point the per-token cost and the data-residency question both evaporate. Qwen 3.7 Max gives you neither. It’s API-only, and that API runs through a China-based provider, the same residency flag GLM’s hosted endpoint carries, except with GLM you have the self-host exit and with Max you don’t. If you push proprietary or regulated code through either hosted API, that one difference is close to the whole decision.

For reference, the Western frontier is still up and to the right, and still expensive. Claude Opus 4.8 leads the hardest coding, around 69 on SWE-bench Pro, a clear few points over both of these, but it charges $25 per million output, and GPT-5.5 sits at $30. Both Chinese models are landing a point or two off the frontier for somewhere between a quarter and a sixth of the price. That, not any single benchmark, is the actual story of 2026.

One more practical line: if you need to look at images, skip both. Qwen 3.7 Max is text-only and so is GLM. Qwen’s own answer there is the cheaper Qwen 3.7 Plus (June 2026, vision and video, $0.40 in and $1.60 out), and otherwise you’re back to Opus.

So which one

Reach for GLM-5.2 when cost rules, when you want to self-host, or when the work is GitHub-issue coding that stays in text. We took it apart in full in our GLM-5.2 breakdown, and it’s still the value pick of the pair.

Reach for Qwen 3.7 Max when the work is heavy on math and reasoning, or when you’re running long autonomous agents that have to stay coherent for hours, and you’re comfortable living on a hosted API with no self-host option.

Stay on Opus 4.8 for the hardest long-horizon coding and anything that touches images, when the budget allows the jump.

The headline that ages best isn’t “Qwen beats GLM,” or the reverse, because on code they don’t. It’s that two models out of Chinese labs are now trading punches a point off the frontier, one of them fully open, both at a fraction of the price. A year ago that reads like a press release. In June 2026 it’s just the leaderboard.

Sources: the GLM-5.2 vs Qwen 3.7 Max head-to-head on CodingFleet, Qwen 3.7 specs and pricing via ofox.ai and Artificial Analysis, plus our own GLM-5.2 figures. Benchmarks and prices as published in June 2026, and this field turns over fast, so re-check before you commit a project.

Frequently asked questions

Is Qwen 3.7 Max better than GLM-5.2 for coding?

On coding they are a tie. GLM-5.2 is a hair ahead on SWE-bench Pro (62.1 vs 60.6) and MCP-Atlas (77.0 vs 76.4), Qwen edges it on HLE without tools (41.4 vs 40.5), and every gap is under 1.5 points, which is inside the run-to-run noise. Pick on cost, openness or reasoning, not on the coding score.

Is Qwen 3.7 Max open-weight or open-source?

No. Qwen 3.7 Max is API-only and proprietary, with no weights to download. The open Qwen models are the smaller 3.6 line, like Qwen3.6-27B. If self-hosting matters to you, GLM-5.2 is the open one here: it ships under an MIT license.

How much does Qwen 3.7 Max cost?

About $2.50 per million input tokens and $7.50 per million output, with cached input near $0.25. That is roughly 1.7x GLM-5.2 on output ($4.40). Qwen also sells a cheaper multimodal tier, Qwen 3.7 Plus, at $0.40 in and $1.60 out.

Does Qwen 3.7 Max support images or vision?

No, Qwen 3.7 Max is text-only, and so is GLM-5.2. If you need image input, Qwen 3.7 Plus (released June 2026) adds vision and costs less, or you go to a frontier model like Claude Opus 4.8.

Qwen 3.7 Max or Claude Opus 4.8?

Opus 4.8 still leads the hardest long-horizon coding, around 69 on SWE-bench Pro against the low 60s for both Chinese models, and it reads images. But it costs $25 per million output against Qwen Max at $7.50. Opus for the ceiling and vision, Qwen Max for value plus math, reasoning and long agent runs.

Tags: aiarticlebenchmarksglmllmqwen
Share199Tweet124
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.