DevNews

Kimi K3 open weights land July 27: the 1.4TB catch

On this page
  1. Do the division before you plan a download
  2. The price list tells you what Moonshot expects you to do
  3. What the demos are actually showing
  4. The license is the part that decides anything
  5. Sources

Somebody on your team is going to ask whether you can just host Kimi K3 yourselves. Short version: the weights land by July 27, and they weigh roughly 1.4 terabytes. That's 2.8 trillion parameters at four bits each, and four bits is already the cheap option. Divide 1.4 TB by an 80 GB accelerator and you need eighteen of them holding nothing but the file, before one token of that 1M context reaches a KV cache. Moonshot doesn't pretend otherwise in its own tech blog, which recommends supernode configurations with 64 or more accelerators. So yes, open. Open the way a container ship is purchasable.

The short answer

Moonshot ships the full Kimi K3 weights by July 27, 2026. The model is 2.8 trillion parameters trained in MXFP4, which puts the file around 1.4 TB and puts self-hosting out of reach for anyone without a rack. The hosted API is already live and priced. The license text, the one thing that decides whether you can ship a product on these weights, isn’t out yet.

1.4 TBweights, MXFP4
64+accelerators, per Moonshot
licensestill unpublished
Answer card: Moonshot releases the full 2.8 trillion parameter Kimi K3 weights by July 27, 2026 in MXFP4 at roughly 1.4 TB, about 18 accelerators of 80 GB before KV cache, with Moonshot recommending supernode configurations of 64 or more accelerators.
The date and the recommendation are both Moonshot's. The 1.4 TB is just division. PNG

Do the division before you plan a download

We wrote about K3 the night it went live in the app, when half the specs floating around were leak-grade. Most of that has since landed in an actual tech blog, including the number that matters here.

2.8 trillion parameters. MXFP4, so half a byte each. Multiply.

You get 1.4 terabytes of weights, and that’s the compressed-by-design version: the same model at 16-bit would be 5.6 TB. Moonshot didn’t quantise after the fact either, it ran quantisation-aware training from the SFT stage onward with MXFP4 weights and MXFP8 activations, which is why four bits doesn’t read as a lossy afterthought here. Nice engineering. It still doesn’t fit anywhere near a workstation.

Log-scale comparison of Kimi K3 weight footprint: 5.6 TB at 16-bit precision, 1.4 TB as shipped in MXFP4, against 0.08 TB for a single 80 GB accelerator.
Log scale, because on a linear one the 80 GB card is a pixel. PNG

Eighteen 80 GB accelerators just to hold the file. That’s the floor, and it’s a floor that does no work: nothing in that 1.4 TB is a KV cache, and K3 advertises a 1M token context window. Which is why Moonshot’s own deployment note is the most useful sentence in the whole blog post. It recommends “supernode configurations with 64 or more accelerators”. Sixty-four, against the eighteen the weights alone need.

Read that gap as the honest hardware spec. Also note MXFP4 wants NVIDIA Blackwell or AMD MI400 class silicon to be native, so an older fleet doesn’t get the format for free.

The price list tells you what Moonshot expects you to do

Almost nobody who “adopts” K3 will download it. They’ll rent it, from Moonshot or from whichever cloud stands it up, and the posted rates are built for that.

Three tiers: $0.30 per million tokens on cache-hit input, $3.00 on cache-miss input, $15.00 on output. That ten times spread between a cache hit and a miss is not decoration. Moonshot reports hit rates above 90 percent in coding workloads, which means a stable prompt prefix is a pricing decision, not a style preference. If your agent rebuilds its system block every call, you’re paying ten times more for the privilege.

For scale: Anthropic lists Claude Opus 5 at $5 in and $25 out per million. So the hosted K3 undercuts a frontier closed model on output by a third, while Moonshot’s own blog puts K3 behind Claude Fable 5 and GPT-5.6 Sol on overall performance. That’s a coherent position, not a contradiction. Cheaper and second.

What the demos are actually showing

The launch material leans hard on one-shot browser builds, and honestly they’re more interesting than the benchmark table nobody can reproduce yet.

Voxel Colosseum browser demo generated by Kimi K3, a Roman amphitheatre filled with crowds inside a blocky city, with an FPS 120 and 39.9k voxel overlay.

Image: Moonshot AI, from the Kimi K3 tech blog.

A voxel colosseum with a crowd in it, running in a browser at the frame counter it prints in the corner. There are eight more like it, plus video captures, all generated rather than hand-built. Reported leaderboard placements point the same way, with K3 debuting top of a frontend coding arena, though we’d flag that the same week produced both a third and a fourth place for K3 on the same independent index, depending on who you read. Snapshots, not standings.

The license is the part that decides anything

Here’s what’s missing, and it isn’t small.

There’s no license text. Moonshot has committed to publishing weights, and coverage expects a modified MIT arrangement because that’s what earlier Kimi models shipped with. The K2.7 Code release carried one attribution clause: show the model name prominently if your product clears roughly 100 million monthly active users or 20 million dollars of monthly revenue. Reasonable, and most teams never touch that threshold.

But expected isn’t published. Until the file lands, you can’t say whether K3 is clean inside a commercial product, under what attribution, or with what limits on training your own model on its outputs. That last one is exactly the clause the open weights lobbying fight is circling right now, so it’s worth reading rather than assuming.

Checklist separating what Moonshot published about Kimi K3, including the July 27 weight date, architecture and pricing, from what is still pending: the license text, an official file size, and consistent independent rankings.
Four things you can plan on, three you can't. The license is the one that blocks shipping. PNG

If you want a model you can genuinely run on your own metal this quarter, K3 isn’t it and was never going to be. Qwen3.7 at the small end is still the honest answer for local work. What K3 gives you is optionality: the weights exist, so if a vendor’s terms turn hostile, somebody can stand this up for you. That’s worth something. Just not a download link.

Sources

The release date, parameter count, expert routing, attention design, MXFP4 with MXFP8 quantisation-aware training, the 64-accelerator deployment recommendation, the cache hit rate and all three price tiers are quoted from Moonshot’s own Kimi K3 tech blog, which also hosts the demo renders and videos. The 1.4 TB and 5.6 TB footprints are our arithmetic on that parameter count, matched by TECHi of July 24, which is also the source for the eighteen-accelerator floor, the roughly 50 billion active parameter equivalent and the point that the license text does not yet exist. The MXFP4 hardware support note and the license-pending status are corroborated by a Hugging Face community write-up. The modified MIT precedent and its attribution threshold come from the earlier K2.7 Code release. Leaderboard placements are as reported by third parties and are not consistent between accounts.

Frequently asked questions

When are the Kimi K3 weights released?

Moonshot states it in its own K3 tech blog: the full model weights will be released by July 27, 2026. The hosted model has been live in the Kimi app and API since the evening of July 16. So the API came first and the download follows about eleven days later, which is the same order Moonshot used for earlier Kimi releases.

How much disk and memory do the Kimi K3 weights need?

About 1.4 TB. Moonshot published the parameter count, 2.8 trillion, and the format, MXFP4, which is half a byte per parameter, so 2.8 trillion times 0.5 bytes lands at roughly 1.4 TB. At 16-bit it would be around 5.6 TB. Moonshot itself does not state a file size anywhere, so treat 1.4 TB as arithmetic rather than a published spec. Weights are also only part of the bill: the KV cache for a 1M token context is extra.

What hardware does it take to actually serve Kimi K3?

More than most people assume. Just holding 1.4 TB of weights takes eighteen accelerators of 80 GB, and Moonshot recommends deploying on supernode configurations with 64 or more accelerators. MXFP4 is natively supported on NVIDIA Blackwell and AMD MI400 class silicon, which is another way of saying older cards will not give you the format for free. Renting inference is the realistic path for almost everybody.

What license do the Kimi K3 weights ship under?

Unpublished as of July 26. Coverage expects a modified MIT license, because that is what earlier Kimi models used, and the K2.7 Code variant carried a single attribution clause requiring products above roughly 100 million monthly active users or 20 million dollars in monthly revenue to display the model name. Expected is not published though. Until the license file lands next to the weights, nobody can tell you cleanly whether K3 is usable in your commercial product or what limits apply to training on its output.

Is the hosted Kimi K3 API cheaper than Claude Opus 5?

On list price, yes. Moonshot posts $0.30 per million tokens for cache-hit input, $3.00 for cache-miss input and $15.00 for output. Anthropic prices Opus 5 at $5 input and $25 output per million in standard mode. The cache tier is the part worth engineering around: Moonshot reports cache hit rates above 90 percent in coding workloads, and a ten times gap between hit and miss means your prompt prefix stability shows up on the invoice.