• Latest
  • Trending
  • All
Official Meta launch artwork for Muse Code and Muse Spark 1.2: a fan of blue lines on a pale background converging into one line that ends inside a small ring.

Meta Muse Code: Muse Spark 1.2 and the 21x cheaper tier

6 August 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Sunday, September 20, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Meta Muse Code: Muse Spark 1.2 and the 21x cheaper tier

by stephane
6 August 2026
in Dev
0
Official Meta launch artwork for Muse Code and Muse Spark 1.2: a fan of blue lines on a pale background converging into one line that ends inside a small ring.
493
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Meta published three coding benchmark charts on Wednesday and finished first on none of them. Odd way to launch, and it tells you what Muse Code is actually selling. It shipped 5 August as a terminal coding agent for macOS and Linux, one curl command to install, running on a new model called Muse Spark 1.2. The closest it gets is Terminal-Bench 2.1, where it takes 82.9 against 86.7 for Claude Opus 5. So the pitch isn't the ceiling. It's the invoice. Meta is offering a contributor tier that drops output tokens from $4.25 per million to $0.20, roughly twenty-one times cheaper, in exchange for permission to train future models on your prompts and completions. Your source code is the currency, and we'd think hard before spending it.

The short answer

Meta released Muse Code in beta on 5 August, a terminal coding agent for macOS and Linux powered by a new Muse Spark 1.2. The engineering is real: persistent background subagents, isolated worktrees, and a crash-safe event log that lets a run resume exactly where it died. The model finishes second or third on every chart Meta published. What it wins is the price, and the cheapest price is paid in your source code.

$0.20output per 1M, contributor tier
82.9Terminal-Bench 2.1, second place
60/minrequest cap on the cheap tier
Official Meta launch artwork for Muse Code and Muse Spark 1.2: a fan of blue lines on a pale background converging into one line that ends inside a small ring. Image: Meta, launch artwork from the Muse Code and Muse Spark 1.2 announcement.

Many threads converging into one. Honestly, the artwork is a better description of the product than the blog post is.

Answer card: Meta released Muse Code in beta on 5 August 2026, a terminal coding agent for macOS and Linux running on Muse Spark 1.2, priced at $1.25 and $4.25 per million tokens on the standard tier or $0.10 and $0.20 on a contributor tier that lets Meta train on your prompts and completions.
A real agent runtime, a mid-pack model, and a discount with a licence attached.

The runtime is the interesting half

Strip the pricing away and there’s genuine engineering here. Muse Code keeps a set of specialised background agents alive for the whole session instead of spawning a fresh one per task, which Meta says cuts the redundant re-reading of the same files that makes long agent runs so slow. Those agents decide for themselves when to report back to the main loop.

Then there’s the event log. Every model call, every tool run, every approval, every edit gets appended to a local log that Meta describes as the single source of truth. The runtime is replay-exact and restart-safe, so a crash three hours into a refactor doesn’t cost you the refactor. The agent picks up precisely where it stopped. If you have ever watched a long agent run die on a network blip and had nothing to show for it, you know why that’s the line in the announcement we’d underline.

Three bundled skills ship by default. /plan turns a task into an approval-gated plan, /grill stress-tests that plan until it holds, and /goal drives toward a stated objective. Approval gating on the plan rather than on each edit is a sensible default, and it’s roughly where the rest of the field has landed too.

The fan-out is the demo everyone quoted. Mark Zuckerberg put it this way on X: “When a job is big enough, it fans out to separate sub-agents working in parallel in isolated worktrees.” Meta says it built six features for a game simultaneously with no collisions in testing. Worth noting that Cursor was swarming agents at a whole SQLite reimplementation back in July, so parallel-agents-in-worktrees is now a pattern rather than a differentiator.

The case study we’d actually chase down is the kernel one. Meta ran the model on GPU kernel optimisation for KDA and MLA on NVIDIA Hopper, over 1,000 tool calls and up to 24 hours per run, writing Triton by hand with third-party kernel libraries explicitly banned so it couldn’t just wrap an existing implementation. Whatever you think of the benchmarks, a 24-hour agent loop that survives to the end is a different class of claim from a 20-minute one.

Meta published charts it loses

This part is unusual, and I’d rather credit it than mock it. Meta put its own coding numbers next to four named rivals and came second or third on all three.

Bar chart of Terminal-Bench 2.1 and DeepSWE 1.1 scores as published by Meta: Muse Spark 1.2 at 82.9 and 59.3, Claude Opus 5 at 86.7 and 65.0, GPT-5.6 Terra at 81.8 and 64.8.
Second on the terminal test, third on the long-horizon one. Meta's own figures.

Terminal-Bench 2.1 has Muse Spark 1.2 at 82.9, behind Claude Opus 5 on Claude Code at 86.7, ahead of GPT-5.6 Terra on Codex at 81.8, Grok 4.5 at 81.6 and Gemini 3.6 Flash at 78.9. DeepSWE 1.1 is worse for Meta: 59.3 against 65.0 and 64.8 for the two leaders. And on Meta’s internal coding bench, the one it controls entirely, it posts 70.6 against 79.4 for Opus 5. Losing on your own private eval is a strange thing to publish and a slightly reassuring one.

One caveat that has been raised and that we can’t resolve: some of the gain from Muse Spark 1.1 to 1.2 reflects the new harness rather than the model. Meta co-trained the two together, which is good product engineering and makes the model number nearly meaningless in isolation. You aren’t buying Muse Spark 1.2. You’re buying Muse Spark 1.2 inside Muse Code.

The discount is a licence, and the cap is the catch

Standard tier runs $1.25 per million input tokens, $4.25 output, $0.15 for cached input, with a no-training commitment attached. That’s the same headline rate as Muse Spark 1.1 in July, so the model got better and the price didn’t move.

The contributor tier is where it gets interesting. $0.10 input, $0.20 output, $0.002 cached. Output is about twenty-one times cheaper and input around twelve. The condition, stated plainly, is that you permit Meta to use your prompts and completions to train future models.

Two-column comparison of the Muse Code standard tier at 1.25 dollars input, 4.25 output, 0.15 cached, 3000 requests per minute and a no-training commitment, against the contributor tier at 0.10 input, 0.20 output, 0.002 cached, 60 requests per minute and training rights over prompts and completions.
The cheap column has two prices on it. Only one of them is in dollars.

Read that through a coding agent’s lens. A prompt here isn’t a question you typed. It’s the file contents the agent pulled in, the diffs it drafted, the stack traces, the config it read on the way past. Whatever your repo contains is what goes into the prompt, which makes this a licensing question for your legal team rather than a line item for your finance team. Meta is accepting zero-data-retention requests separately, and Alexandr Wang described that as “a big enterprise feature that is important for folks”, which reads to us like it is not the default and not free.

Then the rate limits, which nobody seems to be putting next to the feature list. Contributor is capped at 60 requests per minute. Standard gets 3,000 requests and 4 million tokens per minute. Now recall that the marquee capability is fanning a job out to parallel subagents, and every subagent draws requests from that same bucket. Six parallel workers on a 60-per-minute budget is ten requests each per minute. I might be wrong about where that actually binds, since nobody has published a measured run, but the cheap tier and the parallel fan-out look like they’re pulling against each other.

One more practical snag from early testers: the agent refuses to start until billing details are on file, even on the discounted tier. Free trial, this is not.

So who’s it for. If you’re writing open source, or throwaway prototypes, or anything you’d have published anyway, the contributor tier is close to free money and you should try it. If you’re inside a company with a code confidentiality clause, the interesting price is the one you can’t take, and standard tier at $1.25 and $4.25 is a fine rate for a mid-pack model that happens to be wrapped in a crash-safe runtime. Availability got better too, incidentally: Muse Spark 1.1’s API launched as a US-only preview, while Meta says 1.2 ships with expanded global access, though it hasn’t spelled out which regions that covers.

Sources

Product details, the install command, the runtime design and the kernel case study come from Meta Superintelligence Labs, Introducing Muse Code and Muse Spark 1.2, published 5 August 2026. Benchmark scores for all five models are from that announcement as reported by explainX. Tier pricing, rate limits and the training terms are reported by Implicator and BigGo Finance, including the billing-required snag. The Zuckerberg and Wang quotes are via TechCrunch and Engadget.

Frequently asked questions

What is Meta Muse Code?

It is a terminal coding agent Meta Superintelligence Labs released in beta on 5 August 2026, for macOS and Linux, installed with a single curl command from dev.meta.ai. It plans a change, writes the code and validates the result across a large repository, and it can fan work out to persistent subagents running in isolated git worktrees. The model underneath is Muse Spark 1.2, released the same day.

How much does Muse Code cost?

There are two tiers. Standard is $1.25 per million input tokens, $4.25 output and $0.15 for cached input, with a no-training commitment. The contributor tier is $0.10 input, $0.20 output and $0.002 cached, which is about twenty-one times cheaper on output, and it requires you to let Meta train future models on your prompts and completions. Both tiers want billing details on file before the agent will run at all.

Is Muse Spark 1.2 better than Claude Opus 5 or GPT-5.6?

Not on the numbers Meta itself published. Terminal-Bench 2.1 puts Muse Spark 1.2 at 82.9 against 86.7 for Claude Opus 5 and 81.8 for GPT-5.6 Terra. DeepSWE 1.1 puts it at 59.3 against 65.0 and 64.8. On Meta's internal coding bench it scores 70.6 against 79.4 for Opus 5. So it sits mid-pack on frontier coding and wins on price, which is a coherent position, just not the one the launch language implies.

What happens to my code on the contributor tier?

Meta gets permission to use your prompts and completions to train future models. In a coding agent, prompts are not just questions: they carry the file contents the agent read, the diffs it proposed and whatever secrets happen to be sitting in the repo. That is a licensing and compliance decision, not a billing one. Zero data retention is being handled as a separate request process, and Meta's AI chief has described it as an enterprise feature rather than a default.

What is the rate limit on the cheap tier?

60 requests per minute, against 3,000 requests and 4 million tokens per minute on standard. That gap matters more than it looks, because the headline feature is fanning a job out to parallel subagents and every one of those subagents spends requests from the same budget. We have not seen anyone measure where 60 per minute actually binds, so treat it as a thing to test rather than a verdict.

Tags: agentsaillmmetanewspricing
Share197Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.