DevNews

ByteDance's 10T model is pretraining, not a product

On this page
  1. What was actually reported
  2. The total is the memory bill, not the price
  3. China’s biggest is not the world’s biggest
  4. The line about distillation is the actual news
  5. What to do with this today
  6. Sources

Somebody forwarded us the headline twice on Friday. China's biggest AI model, ten trillion parameters. Here's the actual shape of it: on 7 August the Financial Times reported that ByteDance is pretraining a mixture of experts model of up to 10 trillion total parameters, on roughly 30,000 GPUs, over the three to six months a run that size normally takes. Pretraining. Not a launch. The report carries no model name, no chip, no release date, and no active parameter count, which is the one number that would tell you what a token through it might cost. So nobody can price it. What we can do is take the 10 trillion apart, because that figure is carrying the whole headline and it means less than it looks.

The short answer

The Financial Times reported on 7 August 2026 that ByteDance is pretraining a mixture of experts model of up to 10 trillion total parameters on roughly 30,000 GPUs, citing three people familiar with the project. A run that size takes three to six months, so completion lands in 2027 at the earliest. Nothing published names the model, the chips, the release date or the active parameter count. ByteDance has not commented. Treat 10 trillion as a ceiling under consideration, not a spec.

10Ttotal parameters, reported ceiling
0active parameters disclosed
2027earliest the run could finish
Answer card: the Financial Times reported on 7 August 2026 that ByteDance is pretraining a mixture of experts model of up to 10 trillion total parameters on roughly 30,000 GPUs over three to six months, citing three people familiar with the project, with no model name, no active parameter count, no chip named and no release date.
The report in one card. The number in the headline is the least useful thing in it. PNG

What was actually reported

One outlet broke this. The Financial Times, 7 August 2026, sourced to three people familiar with the project. Everything else you have seen since is a relay of that piece, which matters when you are deciding how much weight to put on any single figure in it.

What the FT put on the record: a mixture of experts model, up to 10 trillion total parameters, in pretraining, needing around 30,000 GPUs for a run of three to six months. Zhang Yiming has told the Seed team, roughly 2,000 people, to chase world-leading capability on a long horizon rather than domestic leaderboard position.

What it did not put on the record is longer. No model name. No release date. No chip type, which quietly leaves the entire export-control question open, since the interesting part of assembling 30,000 accelerators in China is not the number but the sourcing. No benchmark, obviously. And no active parameter count.

ByteDance has said nothing either way.

Checklist figure separating what the Financial Times report of 7 August 2026 states on the record about the ByteDance model, including the 10 trillion parameter ceiling, roughly 30,000 GPUs, the 2,000 person Seed team and the instruction to avoid distillation, from what it does not establish, including the model name, the active parameter count, the chips, the release date, any benchmark and any confirmation from ByteDance.
Five things on the record, five things missing. The missing column is where the engineering questions live. PNG

The total is the memory bill, not the price

Here is why we keep circling the active parameter count.

In a mixture of experts model, the total parameter count tells you how much memory the thing occupies, because every weight has to live somewhere whether or not it fires. The active count is the slice actually computed for each token, and that is what drives latency and cost per token. Two numbers, two entirely separate budgets, and headlines only ever carry the first one.

Look at what that gap does in practice. Tencent’s Hy3 is 295B total and 21B active per token. Fourteen to one. A model can be enormous to host and quite cheap to query, and you cannot tell which from the total alone.

Now do it to 10 trillion. At a Hy3-like ratio you would be computing something in the region of 700 billion parameters per token, which is heavy but not absurd for a frontier lab. At a tighter ratio it drops further. At a looser one it becomes ruinous. We genuinely have no idea which, and neither does anybody quoting the headline.

The hosting side is easier to reason about, and grim. Kimi K3 at 2.8 trillion parameters lands at roughly 1.4 TB in four-bit format. Scale that arithmetic to 10 trillion and you are looking at something like five terabytes of weights at the same precision, before a single token of context touches a KV cache. That is not a self-hosting story for anyone. It is barely a story for most clouds.

China’s biggest is not the world’s biggest

The comparison everyone wants is the one the evidence supports least.

Ten trillion would make this comfortably the largest Chinese model, more than three times Kimi K3. Fine. But the FT-relaying coverage then sets it against Anthropic’s Mythos 5 at “about 8 trillion parameters”, and that figure is an industry estimate. Anthropic has never published a parameter count for Mythos. OpenAI publishes nothing of the kind either. The Decoder’s write-up notes Musk has referred to Grok variants at 6 and 10 trillion parameters, training on Colossus 2.

So the ranking being drawn is: a leaked ceiling, versus an outside guess, versus a founder’s remarks. I would not bet a serving decision on any of it.

Bar comparison of total parameter counts showing the ByteDance model at a reported ceiling of about 10 trillion, Anthropic Mythos 5 at an industry estimate of about 8 trillion which Anthropic has never disclosed, and Moonshot Kimi K3 at a published 2.8 trillion, with a note explaining that only the Kimi K3 figure came from the lab that built the model.
Three bars, three completely different standards of evidence. Only the short one was published by the lab that built it. PNG

Scale also stopped being a clean proxy for capability a while ago. Data quality and training technique move results as much as raw count does, which is why a 2.8 trillion parameter Chinese model already trades benchmarks with things far larger. Ten trillion parameters trained badly is an expensive way to be mid.

The line about distillation is the actual news

Buried under the big number sits the part we think matters more.

Zhang told an internal meeting that ByteDance will not distill from competitors’ outputs, even if holding that line costs the company short-term standing against domestic rivals. Reporting says it has stuck to that for over a year.

That is a real strategic commitment with a real price attached, and it says something about the next twelve months of Chinese open weights. Distillation lineage is exactly what makes provenance and licensing awkward when you pull a set of weights into a product. A lab publicly refusing the shortcut is a lab whose training story you can more plausibly reason about later. Whether it produces a better model is a separate question, and honestly I would guess it costs them a few months against rivals who keep distilling.

Zhang is also framing this as an eighteen-month-plus project rather than a quarter, which fits the arithmetic. Thirty thousand GPUs, three to six months of pretraining, then post-training and evaluation on top. Nothing about this reaches an API in 2026.

What to do with this today

Nothing, mostly. That is the honest answer, and it is fine.

If you are choosing a model this month, this story has no input for you. There is no endpoint, no price, no eval. Keep testing what is actually callable.

If you are planning around Chinese open weights for next year, note two things instead of the big number. First, ByteDance has never open-weighted a frontier Seed model, and nothing in this report says it intends to, so a 10 trillion parameter Doubao successor is not obviously a thing you will ever download. Second, the distillation stance is a provenance signal worth tracking, more than the scale is.

And when the model does land, the first question to ask is not how many parameters. It is how many are active.

Sources

Original reporting: the Financial Times piece of 7 August 2026, citing three people familiar with the project, which sits behind a paywall. Relays we read and checked against each other: The Next Web, 7 August, for the parameter ceiling and the Kimi K3 comparison. The Decoder, 7 August, for the Mythos 5 figure explicitly qualified as an industry estimate, the Grok parameter remarks and the Seed team instruction. Crypto Briefing, 8 August, for the 30,000 GPU count, the mixture of experts architecture and the Seed team size. Slashdot’s summary for the caveat that parameter count is not everything. The Kimi K3 size and the four-bit file arithmetic are Moonshot’s own published figures, worked through in our Kimi K3 weights piece. Nothing here is confirmed by ByteDance.

Frequently asked questions

Has ByteDance released a 10 trillion parameter model?

No. The Financial Times reported on 7 August 2026 that a model of up to 10 trillion total parameters is in pretraining, citing three people familiar with the project. Pretraining is the phase before anything is fine-tuned, evaluated or shipped, and it typically runs three to six months on a cluster this size. There is no model name, no release date and no benchmark, because there is no finished model. ByteDance has neither confirmed nor denied the report.

Is 10 trillion parameters bigger than GPT or Claude?

Nobody can answer that cleanly, and that is the point. Anthropic has never published a parameter count for Mythos 5; the roughly 8 trillion figure you see quoted is an outside industry estimate. OpenAI does not publish counts either. The Decoder notes that Elon Musk has talked about Grok variants at 6 and 10 trillion parameters. So 10 trillion would make this the largest Chinese model by a wide margin, and comparing it to Western frontier models means comparing a leak against a guess.

Why does the active parameter count matter more than the total?

Because in a mixture of experts model they pay for different things. Total parameters set your memory bill, since every weight has to sit somewhere. Active parameters are the slice actually computed per token, and that sets your latency and your cost per token. Tencent's Hy3 is 295B total but only 21B active, roughly a fourteen to one ratio. Apply anything like that to 10 trillion and the compute per token could look ordinary while the memory footprint stays enormous. The FT report gives no active count, so both halves of the economics are unknown.

How does this compare to Kimi K3?

Kimi K3 is the biggest Chinese model whose size its own lab published: 2.8 trillion parameters, which in MXFP4 is a file of roughly 1.4 TB. A 10 trillion parameter model would be more than three times that. Worth noting that coverage relaying the FT quotes Kimi K3 at both 2.8 trillion and about 3.3 trillion, so even the comparison baseline wobbles depending on which write-up you read.

What is the distillation detail everyone skipped?

Zhang Yiming told an internal meeting that ByteDance will not use distillation on competitors' outputs, even if that costs it short-term standing against domestic rivals, and reporting says the company has held that line for over a year. He instructed the Seed team, around 2,000 people, to aim for world-leading capability over the long term. For anyone consuming Chinese open weights that is arguably the more useful signal in the story, since distillation lineage is what makes license and provenance questions messy.