• Latest
  • Trending
  • All
Writer product screenshot from the Palmyra X6 announcement, showing a model picker with Palmyra X6 selected above Palmyra X5, next to an AI visibility audit playbook dashboard listing agent runs with status, latency in seconds, average cost and total spend per row.

Palmyra X6 runs on GLM-5.2, priced $2 in and $8 out

15 August 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Monday, September 21, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Palmyra X6 runs on GLM-5.2, priced $2 in and $8 out

by stephane
15 August 2026
in Dev
0
Writer product screenshot from the Palmyra X6 announcement, showing a model picker with Palmyra X6 selected above Palmyra X5, next to an AI visibility audit playbook dashboard listing agent runs with status, latency in seconds, average cost and total spend per row.
492
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Read the fourth paragraph of the press release and the interesting sentence is right there, unhedged: Palmyra X6 was post-trained on top of GLM-5.2, which Writer calls the strongest available open weight model. So the new enterprise flagship from a San Francisco company backed by Salesforce and Adobe and IBM is a Chinese open weight model with a finishing pass on it, and Writer prints that fact rather than burying it. It shipped on 13 August at $2 per million input tokens and $8 per million output, with a headline score of 0.87 that beats Claude Opus 4.8 by a hundredth of a point. That score is worth about ten minutes of your attention, and not for the reason the headline suggests.

The short answer

Writer shipped Palmyra X6 on 13 August, post-trained on Z.ai’s open weight GLM-5.2, at $2 per million input and $8 per million output. It claims 0.87 out of 1.00 against Claude Opus 4.8’s 0.86 on nine internal evaluations of enterprise marketing workflows. The score is vendor run and narrow by design. The harness result underneath it is the part worth stealing.

GLM-5.2the base Writer post-trained
$2 / $8per million in and out
0.87on Writer own nine evals
Answer card: Writer released Palmyra X6 on 13 August 2026, post-trained on top of the open weight GLM-5.2 from Z.ai, priced at 2 dollars per million input tokens and 8 dollars per million output tokens, with a headline score of 0.87 out of 1.00 from Writer's own nine internal enterprise evaluations.
The base model is named in the release. That is rarer than it should be.

What shipped

Three things on 13 August, and the model is only one of them.

Palmyra X6, the new flagship, aimed squarely at marketing and revenue workflows. A rebuilt Writer Agent harness, which is the orchestration layer that plans a task, calls tools and hands work to sub-agents. And a governance console that shows admins which agent workflows are burning tokens, with alerts and spending limits attached.

That last one is the least glamorous and probably the reason anyone signs the renewal. Writer quotes IDC saying three quarters of organisations now name excessive AI spending as a major risk to their plans, which sounds like vendor framing until you have watched an agent loop for four minutes on a task a human would have finished in one.

Writer's Palmyra X6 announcement image: a model picker showing Palmyra X6 selected above Palmyra X5, beside a dashboard titled AI visibility audit playbook with columns for status, latency in seconds, average cost and total spend for each agent run.

Image: WRITER

Look at what the announcement image is actually showing. A model picker on the left, and on the right a spend dashboard with a dollar column per row. Not a benchmark chart. That is the pitch.

The base model is the story

X6 was post-trained on top of GLM-5.2.

Writer states it directly, in the paragraph where a vendor would normally reach for the word “proprietary”. We covered GLM-5.2 when it landed as a 753 billion parameter mixture of experts with only eight experts active per token, and GLM-5.3 arrived a couple of days ago with the weights promised on a two week delay. So the lineage here is public, and the base is downloadable by anyone.

What Writer sells on top is the post-training recipe, the harness, the governance layer and an enterprise contract with a US company. Honestly, that is a defensible business. It is also a quietly enormous signal about where the open weight frontier sits in August 2026: an enterprise vendor with Fortune 500 logos on the wall looked at everything available, picked a model out of Beijing as its foundation, and said so in a press release.

If your procurement process has a question about model provenance, this is the release that will surface it. Better you find that out from a paragraph than from a security review three months into a rollout.

Bar chart of blended cost per million tokens at a three to one input to output mix, computed from the list prices Writer published: Palmyra X6 at $3.50, Gemini 3.1 at $4.38, Claude Sonnet 4.6 at $6.00, GPT-5.5 at $7.50 and Claude Opus 4.8 at $30.00, with each model's score on Writer's nine internal evaluations shown alongside.
Four of five models within a tenth of a point. Nine to one on price.

Read the 0.87 properly

Writer built its own evaluation suite. It says so, it explains why, and it describes what is in it: nine capabilities including grounding and retrieval, tool use, content generation, sub-agent delegation and brand voice, drawn from workflows its customers actually run.

X6 scored 0.87 out of 1.00. Claude Opus 4.8 got 0.86 at $15 and $75 per million. Claude Sonnet 4.6 got 0.85 at $3 and $15. GPT-5.5 got 0.80 at $5 and $15. Gemini 3.1 got 0.77 at $2.50 and $10.

Now, the honest reading. This is not a leaderboard result and Writer never claims it is. It is a vendor measuring five models on its own workloads, inside its own orchestration harness, with one of the five tuned specifically to run in that harness. Of course it wins. The finding is not “X6 beats Opus”, it is “for this narrow class of work, the cheap tuned model was good enough that the expensive one stopped being worth nine times the money”.

Which is a real finding. I just would not carry that 0.87 into any sentence about general capability, and neither should the deck someone builds from it.

The cluster is the tell. Four of the five models sit within 0.10 of each other while the blended prices spread from $3.50 to $30. When a benchmark stops separating models, it has usually stopped measuring the thing you are choosing between.

The number we would actually steal

Buried under the model launch is the part that applies whether or not you ever open a Writer contract.

Writer published research on its own orchestration layer, and reports that the rebuilt harness completed tasks 44 percent faster at 41 percent lower cost per task across every model it tested, including third party ones. Not just its own. The mechanism it describes is unglamorous: adapt reasoning depth to the task instead of planning everything, batch high volume work, delegate to sub-agents so context stops getting resent.

Paired specifically with X6 the figures are 52 percent lower cost, 48 percent faster, 10 percent higher quality.

We have seen the same shape in our own agent work at a much smaller scale. The token bill is set by how many times you resend context, not by which model reads it. Swapping a frontier model for a cheap one saves you maybe half. Fixing a harness that re-sends the full history on every tool call saves you more than that, and it works on whichever model you switch to next.

One caveat on the per task dollar figures floating around. Coverage puts the new cost per task near $0.12, down from either $0.21 or $0.25 depending on which write up you read. Writer’s own release does not print a per task figure at all, so treat the specific dollars as reporting rather than as published data.

Checklist separating what Writer confirmed in its own Palmyra X6 press release, including the GLM-5.2 base, the $2 and $8 pricing, the nine evaluation capabilities and the speed figures, from what is missing, including public benchmark results, any parameter count for X6, a consistent cost per task figure and any weights release.
Five things printed by the vendor, four that are not.

Would we use it

Depends entirely on what you are.

If you run marketing or revenue operations at a company with a compliance function, this is a straightforward pitch, and the governance console is the reason rather than the model. Per workflow spend visibility with limits attached is the thing most agent deployments are missing right now, and it is why pilots stall at pilot.

If you are a developer building your own agents, X6 is not for you and Writer is not pretending otherwise. There is no weights release and no self hosting path. The base is right there though, and GLM-5.2 is downloadable today. You would be doing your own post-training and your own harness, which is exactly the work Writer is charging for.

And if you are just tracking prices, the useful data point is that a vendor put $2 and $8 next to Opus 4.8’s $15 and $75 in its own comparison table and expected that to read as reasonable. Six months ago the flagship premium was assumed. It is now something you have to justify per workload.

X6 completes tasks in 26 seconds on average, generates 82 tokens per second, and Writer says it can work unattended toward a single goal for up to eight hours. That last claim is the one I would want to test before believing. Eight hours of coherent unattended agent work is a strong statement, and nothing in the release explains what “sustaining coherent reasoning” was measured against.

Sources

WRITER, WRITER Makes Agentic AI Economically Sustainable at Enterprise Scale With Palmyra X6 Release and Major Harness Upgrades, for the 13 August 2026 release date, the statement that X6 was post-trained on top of GLM-5.2, the $2 and $8 pricing, the 0.87 score and the list prices and scores for Claude Opus 4.8, Claude Sonnet 4.6, GPT-5.5 and Gemini 3.1, the nine evaluated capabilities, the 26 second average task, 82 tokens per second and eight hour unattended figures, the 52 percent, 48 percent and 10 percent harness numbers with X6, the 44 percent and 41 percent figures across all models tested, the IDC quote on AI spending as a risk, and the governance and consumption control features. TechCrunch, Writer introduces new AI model and upgraded harness to contain token costs, for independent confirmation of the GLM-5.2 base, the availability date for Writer customers, and the model agnostic positioning. SiliconANGLE, Writer launches major agentic AI improvements with Palmyra X6 flagship model, and CMSWire, WRITER Releases Palmyra X6, Upgrades Enterprise AI Agent Platform, for the Palmyra X5 lineage, the cost per task reporting and the multi model support through Azure, Bedrock and NVIDIA NIM. Announcement image by WRITER.

Frequently asked questions

What model is Palmyra X6 built on?

GLM-5.2, the open weight mixture of experts model from Z.ai in Beijing. Writer says so in its own release: X6 was post-trained on top of GLM-5.2, which it describes as the strongest available open weight model. Writer does not publish a parameter count for X6 itself, so the only public size figures belong to the base.

How much does Palmyra X6 cost?

$2 per million input tokens and $8 per million output tokens. At a 3 to 1 input to output mix that blends to about $3.50 per million, which is our arithmetic rather than a Writer figure. For comparison, the list prices Writer printed next to its own were $15 and $75 for Claude Opus 4.8, $3 and $15 for Claude Sonnet 4.6, $5 and $15 for GPT-5.5, and $2.50 and $10 for Gemini 3.1.

Is the 0.87 score a public benchmark?

No, and Writer is upfront about that. The suite is nine internal evaluations built from customer production workflows, covering grounding and retrieval, tool use, content generation, sub-agent delegation and brand voice. It measures fitness for marketing and revenue work inside the Writer Agent harness. You cannot line those numbers up against any public leaderboard.

Can I download or self host Palmyra X6?

No. X6 is available to Writer customers on the Writer platform, and there is no weights release. The base model is a different story: GLM-5.2 is open weight, so the thing underneath X6 is downloadable even though X6 is not.

What did the harness upgrade actually change?

Writer rebuilt the orchestration layer to adapt reasoning depth to the task, batch work and delegate to sub-agents. It reports 44 percent faster completion at 41 percent lower cost per task across every model it tested, its own and third party alike. Paired specifically with X6 the figures are 52 percent lower cost, 48 percent faster and 10 percent better quality.

Tags: agentsaienterprisellmnewsopen-weights
Share197Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.