• Latest
  • Trending
  • All
Answer card: the Financial Times reported on 7 August 2026 that ByteDance is pretraining a mixture of experts model of up to 10 trillion total parameters on roughly 30,000 GPUs over three to six months, with no model name, no active parameter count, no chip and no release date disclosed.

ByteDance’s 10T model is pretraining, not a product

9 August 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Monday, September 21, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

ByteDance’s 10T model is pretraining, not a product

by stephane
9 August 2026
in Dev
0
Answer card: the Financial Times reported on 7 August 2026 that ByteDance is pretraining a mixture of experts model of up to 10 trillion total parameters on roughly 30,000 GPUs over three to six months, with no model name, no active parameter count, no chip and no release date disclosed.
494
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Somebody forwarded us the headline twice on Friday. China's biggest AI model, ten trillion parameters. Here's the actual shape of it: on 7 August the Financial Times reported that ByteDance is pretraining a mixture of experts model of up to 10 trillion total parameters, on roughly 30,000 GPUs, over the three to six months a run that size normally takes. Pretraining. Not a launch. The report carries no model name, no chip, no release date, and no active parameter count, which is the one number that would tell you what a token through it might cost. So nobody can price it. What we can do is take the 10 trillion apart, because that figure is carrying the whole headline and it means less than it looks.

The short answer

The Financial Times reported on 7 August 2026 that ByteDance is pretraining a mixture of experts model of up to 10 trillion total parameters on roughly 30,000 GPUs, citing three people familiar with the project. A run that size takes three to six months, so completion lands in 2027 at the earliest. Nothing published names the model, the chips, the release date or the active parameter count. ByteDance has not commented. Treat 10 trillion as a ceiling under consideration, not a spec.

10Ttotal parameters, reported ceiling
0active parameters disclosed
2027earliest the run could finish
Answer card: the Financial Times reported on 7 August 2026 that ByteDance is pretraining a mixture of experts model of up to 10 trillion total parameters on roughly 30,000 GPUs over three to six months, citing three people familiar with the project, with no model name, no active parameter count, no chip named and no release date.
The report in one card. The number in the headline is the least useful thing in it.

What was actually reported

One outlet broke this. The Financial Times, 7 August 2026, sourced to three people familiar with the project. Everything else you have seen since is a relay of that piece, which matters when you are deciding how much weight to put on any single figure in it.

What the FT put on the record: a mixture of experts model, up to 10 trillion total parameters, in pretraining, needing around 30,000 GPUs for a run of three to six months. Zhang Yiming has told the Seed team, roughly 2,000 people, to chase world-leading capability on a long horizon rather than domestic leaderboard position.

What it did not put on the record is longer. No model name. No release date. No chip type, which quietly leaves the entire export-control question open, since the interesting part of assembling 30,000 accelerators in China is not the number but the sourcing. No benchmark, obviously. And no active parameter count.

ByteDance has said nothing either way.

Checklist figure separating what the Financial Times report of 7 August 2026 states on the record about the ByteDance model, including the 10 trillion parameter ceiling, roughly 30,000 GPUs, the 2,000 person Seed team and the instruction to avoid distillation, from what it does not establish, including the model name, the active parameter count, the chips, the release date, any benchmark and any confirmation from ByteDance.
Five things on the record, five things missing. The missing column is where the engineering questions live.

The total is the memory bill, not the price

Here is why we keep circling the active parameter count.

In a mixture of experts model, the total parameter count tells you how much memory the thing occupies, because every weight has to live somewhere whether or not it fires. The active count is the slice actually computed for each token, and that is what drives latency and cost per token. Two numbers, two entirely separate budgets, and headlines only ever carry the first one.

Look at what that gap does in practice. Tencent’s Hy3 is 295B total and 21B active per token. Fourteen to one. A model can be enormous to host and quite cheap to query, and you cannot tell which from the total alone.

Now do it to 10 trillion. At a Hy3-like ratio you would be computing something in the region of 700 billion parameters per token, which is heavy but not absurd for a frontier lab. At a tighter ratio it drops further. At a looser one it becomes ruinous. We genuinely have no idea which, and neither does anybody quoting the headline.

The hosting side is easier to reason about, and grim. Kimi K3 at 2.8 trillion parameters lands at roughly 1.4 TB in four-bit format. Scale that arithmetic to 10 trillion and you are looking at something like five terabytes of weights at the same precision, before a single token of context touches a KV cache. That is not a self-hosting story for anyone. It is barely a story for most clouds.

China’s biggest is not the world’s biggest

The comparison everyone wants is the one the evidence supports least.

Ten trillion would make this comfortably the largest Chinese model, more than three times Kimi K3. Fine. But the FT-relaying coverage then sets it against Anthropic’s Mythos 5 at “about 8 trillion parameters”, and that figure is an industry estimate. Anthropic has never published a parameter count for Mythos. OpenAI publishes nothing of the kind either. The Decoder’s write-up notes Musk has referred to Grok variants at 6 and 10 trillion parameters, training on Colossus 2.

So the ranking being drawn is: a leaked ceiling, versus an outside guess, versus a founder’s remarks. I would not bet a serving decision on any of it.

Bar comparison of total parameter counts showing the ByteDance model at a reported ceiling of about 10 trillion, Anthropic Mythos 5 at an industry estimate of about 8 trillion which Anthropic has never disclosed, and Moonshot Kimi K3 at a published 2.8 trillion, with a note explaining that only the Kimi K3 figure came from the lab that built the model.
Three bars, three completely different standards of evidence. Only the short one was published by the lab that built it.

Scale also stopped being a clean proxy for capability a while ago. Data quality and training technique move results as much as raw count does, which is why a 2.8 trillion parameter Chinese model already trades benchmarks with things far larger. Ten trillion parameters trained badly is an expensive way to be mid.

The line about distillation is the actual news

Buried under the big number sits the part we think matters more.

Zhang told an internal meeting that ByteDance will not distill from competitors’ outputs, even if holding that line costs the company short-term standing against domestic rivals. Reporting says it has stuck to that for over a year.

That is a real strategic commitment with a real price attached, and it says something about the next twelve months of Chinese open weights. Distillation lineage is exactly what makes provenance and licensing awkward when you pull a set of weights into a product. A lab publicly refusing the shortcut is a lab whose training story you can more plausibly reason about later. Whether it produces a better model is a separate question, and honestly I would guess it costs them a few months against rivals who keep distilling.

Zhang is also framing this as an eighteen-month-plus project rather than a quarter, which fits the arithmetic. Thirty thousand GPUs, three to six months of pretraining, then post-training and evaluation on top. Nothing about this reaches an API in 2026.

What to do with this today

Nothing, mostly. That is the honest answer, and it is fine.

If you are choosing a model this month, this story has no input for you. There is no endpoint, no price, no eval. Keep testing what is actually callable.

If you are planning around Chinese open weights for next year, note two things instead of the big number. First, ByteDance has never open-weighted a frontier Seed model, and nothing in this report says it intends to, so a 10 trillion parameter Doubao successor is not obviously a thing you will ever download. Second, the distillation stance is a provenance signal worth tracking, more than the scale is.

And when the model does land, the first question to ask is not how many parameters. It is how many are active.

Sources

Original reporting: the Financial Times piece of 7 August 2026, citing three people familiar with the project, which sits behind a paywall. Relays we read and checked against each other: The Next Web, 7 August, for the parameter ceiling and the Kimi K3 comparison. The Decoder, 7 August, for the Mythos 5 figure explicitly qualified as an industry estimate, the Grok parameter remarks and the Seed team instruction. Crypto Briefing, 8 August, for the 30,000 GPU count, the mixture of experts architecture and the Seed team size. Slashdot’s summary for the caveat that parameter count is not everything. The Kimi K3 size and the four-bit file arithmetic are Moonshot’s own published figures, worked through in our Kimi K3 weights piece. Nothing here is confirmed by ByteDance.

Frequently asked questions

Has ByteDance released a 10 trillion parameter model?

No. The Financial Times reported on 7 August 2026 that a model of up to 10 trillion total parameters is in pretraining, citing three people familiar with the project. Pretraining is the phase before anything is fine-tuned, evaluated or shipped, and it typically runs three to six months on a cluster this size. There is no model name, no release date and no benchmark, because there is no finished model. ByteDance has neither confirmed nor denied the report.

Is 10 trillion parameters bigger than GPT or Claude?

Nobody can answer that cleanly, and that is the point. Anthropic has never published a parameter count for Mythos 5; the roughly 8 trillion figure you see quoted is an outside industry estimate. OpenAI does not publish counts either. The Decoder notes that Elon Musk has talked about Grok variants at 6 and 10 trillion parameters. So 10 trillion would make this the largest Chinese model by a wide margin, and comparing it to Western frontier models means comparing a leak against a guess.

Why does the active parameter count matter more than the total?

Because in a mixture of experts model they pay for different things. Total parameters set your memory bill, since every weight has to sit somewhere. Active parameters are the slice actually computed per token, and that sets your latency and your cost per token. Tencent's Hy3 is 295B total but only 21B active, roughly a fourteen to one ratio. Apply anything like that to 10 trillion and the compute per token could look ordinary while the memory footprint stays enormous. The FT report gives no active count, so both halves of the economics are unknown.

How does this compare to Kimi K3?

Kimi K3 is the biggest Chinese model whose size its own lab published: 2.8 trillion parameters, which in MXFP4 is a file of roughly 1.4 TB. A 10 trillion parameter model would be more than three times that. Worth noting that coverage relaying the FT quotes Kimi K3 at both 2.8 trillion and about 3.3 trillion, so even the comparison baseline wobbles depending on which write-up you read.

What is the distillation detail everyone skipped?

Zhang Yiming told an internal meeting that ByteDance will not use distillation on competitors' outputs, even if that costs it short-term standing against domestic rivals, and reporting says the company has held that line for over a year. He instructed the Seed team, around 2,000 people, to aim for world-leading capability over the long term. For anyone consuming Chinese open weights that is arguably the more useful signal in the story, since distillation lineage is what makes license and provenance questions messy.

Tags: aibytedancechinallmmoenews
Share198Tweet124
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.