• Latest
  • Trending
  • All
Voxel Colosseum scene rendered in a browser demo generated by Kimi K3, showing a Roman amphitheatre packed with crowds inside a blocky city, with an FPS 120 overlay in the corner.

The Kimi K3 weights are 1.4 TB, and that is the catch

3 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

7 September 2026
Answer card stating that Mullvad announced on 3 September 2026 that it is shutting down its public encrypted domain name system servers on 2 November 2026 and sponsoring the Quad9 Foundation instead, with 194.242.2.2 and its five sibling addresses all going away, and virtual private network customers unaffected.

Mullvad’s DNS servers go dark on 2 November, and Quad9 blocks no ads

5 September 2026
OpenAI announcement image for GPT-6 Astra, a spiral galaxy of white, blue and amber points of light curling around a bright core on a near black star field.

GPT-6 Astra lists at $10 and $50, 2.5x what GPT-5.6 Sol costs

6 September 2026
Google's official announcement image for the release, reading Introducing Gemini 3.8 Flash and 3.8 Flash Cyber in black type over a pale blue background with a blurred white chevron and the four colour Gemini spark below.

Gemini 3.8 Flash keeps the price and the 1 January cliff

3 September 2026
Answer card stating that Anthropic announced Enterprise Frontier Safeguards on 1 September 2026, that activity data used for misuse monitoring moves into cloud storage the customer controls under the customer own encryption keys, that Anthropic charges nothing for the feature while the cloud provider bills storage and egress, and that the phased rollout starts later in autumn 2026 with interim zero data retention on Fable 5 and Fable 5.1 for eligible customers.

Anthropic moves retention into your own cloud, for 30 days

3 September 2026
Official Google diagram of a client connection in three numbered steps: a DNS lookup with a query and an address, a TLS ClientHello and ServerHello, then a content exchange with a website. A callout on the DNS step reads 25% of global web traffic is now protected by encrypted DNS, and a callout beside an Android phone on the ClientHello step reads Android 17 supports ECH GREASE by default.

Android 17 hides the SNI, not your DNS or destination

3 September 2026
Still frame from the Claude Fable 5.1 launch video showing model-designed protein binders in orange docked against twelve grey target proteins, rendered as ESMFold2 structure predictions.

Claude Fable 5.1 breaks forced tool use, cuts cache 75%

1 September 2026
Answer card stating that on 31 August 2026 the European Commission designated ChatGPT a Very Large Online Search Engine under the Digital Services Act, the first conversational AI service classified that way, because it answers user prompts and queries including by searching the web, with OpenAI having declared roughly 159.1 million average monthly users in the European Union for ChatGPT search.

The EU now calls ChatGPT a very large search engine

3 September 2026
Answer card stating that on 31 August 2026 the Department of War added OpenAI ChatGPT Mil and Starshield AI Grok for Government to the GenAI.mil portal alongside Google Gemini, all three accredited at Impact Level 5 for Controlled Unclassified Information, with 1.7 million unique users onboarded out of roughly 3 million eligible personnel, and ChatGPT Mil currently serving GPT-5.4 Terra with GPT-5.6 Terra said to be rolling out.

ChatGPT Mil and Grok reached IL5 on GenAI.mil

3 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Friday, September 11, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

The Kimi K3 weights are 1.4 TB, and that is the catch

by stephane
3 September 2026
in Dev
0
Voxel Colosseum scene rendered in a browser demo generated by Kimi K3, showing a Roman amphitheatre packed with crowds inside a blocky city, with an FPS 120 overlay in the corner.
493
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Somebody on your team is going to ask whether you can just host Kimi K3 yourselves. Short version: the weights land by July 27, and they weigh roughly 1.4 terabytes. That's 2.8 trillion parameters at four bits each, and four bits is already the cheap option. Divide 1.4 TB by an 80 GB accelerator and you need eighteen of them holding nothing but the file, before one token of that 1M context reaches a KV cache. Moonshot doesn't pretend otherwise in its own tech blog, which recommends supernode configurations with 64 or more accelerators. So yes, open. Open the way a container ship is purchasable.

The short answer

Moonshot ships the full Kimi K3 weights by July 27, 2026. The model is 2.8 trillion parameters trained in MXFP4, which puts the file around 1.4 TB and puts self-hosting out of reach for anyone without a rack. The hosted API is already live and priced. The license text, the one thing that decides whether you can ship a product on these weights, isn’t out yet.

1.4 TBweights, MXFP4
64+accelerators, per Moonshot
licensestill unpublished
Answer card: Moonshot releases the full 2.8 trillion parameter Kimi K3 weights by July 27, 2026 in MXFP4 at roughly 1.4 TB, about 18 accelerators of 80 GB before KV cache, with Moonshot recommending supernode configurations of 64 or more accelerators.
The date and the recommendation are both Moonshot's. The 1.4 TB is just division.

Do the division before you plan a download

We wrote about K3 the night it went live in the app, when half the specs floating around were leak-grade. Most of that has since landed in an actual tech blog, including the number that matters here.

2.8 trillion parameters. MXFP4, so half a byte each. Multiply.

You get 1.4 terabytes of weights, and that’s the compressed-by-design version: the same model at 16-bit would be 5.6 TB. Moonshot didn’t quantise after the fact either, it ran quantisation-aware training from the SFT stage onward with MXFP4 weights and MXFP8 activations, which is why four bits doesn’t read as a lossy afterthought here. Nice engineering. It still doesn’t fit anywhere near a workstation.

Log-scale comparison of Kimi K3 weight footprint: 5.6 TB at 16-bit precision, 1.4 TB as shipped in MXFP4, against 0.08 TB for a single 80 GB accelerator.
Log scale, because on a linear one the 80 GB card is a pixel.

Eighteen 80 GB accelerators just to hold the file. That’s the floor, and it’s a floor that does no work: nothing in that 1.4 TB is a KV cache, and K3 advertises a 1M token context window. Which is why Moonshot’s own deployment note is the most useful sentence in the whole blog post. It recommends “supernode configurations with 64 or more accelerators”. Sixty-four, against the eighteen the weights alone need.

Read that gap as the honest hardware spec. Also note MXFP4 wants NVIDIA Blackwell or AMD MI400 class silicon to be native, so an older fleet doesn’t get the format for free.

The price list tells you what Moonshot expects you to do

Almost nobody who “adopts” K3 will download it. They’ll rent it, from Moonshot or from whichever cloud stands it up, and the posted rates are built for that.

Three tiers: $0.30 per million tokens on cache-hit input, $3.00 on cache-miss input, $15.00 on output. That ten times spread between a cache hit and a miss is not decoration. Moonshot reports hit rates above 90 percent in coding workloads, which means a stable prompt prefix is a pricing decision, not a style preference. If your agent rebuilds its system block every call, you’re paying ten times more for the privilege.

For scale: Anthropic lists Claude Opus 5 at $5 in and $25 out per million. So the hosted K3 undercuts a frontier closed model on output by a third, while Moonshot’s own blog puts K3 behind Claude Fable 5 and GPT-5.6 Sol on overall performance. That’s a coherent position, not a contradiction. Cheaper and second.

What the demos are actually showing

The launch material leans hard on one-shot browser builds, and honestly they’re more interesting than the benchmark table nobody can reproduce yet.

Voxel Colosseum browser demo generated by Kimi K3, a Roman amphitheatre filled with crowds inside a blocky city, with an FPS 120 and 39.9k voxel overlay.

Image: Moonshot AI, from the Kimi K3 tech blog.

A voxel colosseum with a crowd in it, running in a browser at the frame counter it prints in the corner. There are eight more like it, plus video captures, all generated rather than hand-built. Reported leaderboard placements point the same way, with K3 debuting top of a frontend coding arena, though we’d flag that the same week produced both a third and a fourth place for K3 on the same independent index, depending on who you read. Snapshots, not standings.

The license is the part that decides anything

Here’s what’s missing, and it isn’t small.

There’s no license text. Moonshot has committed to publishing weights, and coverage expects a modified MIT arrangement because that’s what earlier Kimi models shipped with. The K2.7 Code release carried one attribution clause: show the model name prominently if your product clears roughly 100 million monthly active users or 20 million dollars of monthly revenue. Reasonable, and most teams never touch that threshold.

But expected isn’t published. Until the file lands, you can’t say whether K3 is clean inside a commercial product, under what attribution, or with what limits on training your own model on its outputs. That last one is exactly the clause the open weights lobbying fight is circling right now, so it’s worth reading rather than assuming.

Checklist separating what Moonshot published about Kimi K3, including the July 27 weight date, architecture and pricing, from what is still pending: the license text, an official file size, and consistent independent rankings.
Four things you can plan on, three you can't. The license is the one that blocks shipping.

If you want a model you can genuinely run on your own metal this quarter, K3 isn’t it and was never going to be. Qwen3.7 at the small end is still the honest answer for local work. What K3 gives you is optionality: the weights exist, so if a vendor’s terms turn hostile, somebody can stand this up for you. That’s worth something. Just not a download link.

Sources

The release date, parameter count, expert routing, attention design, MXFP4 with MXFP8 quantisation-aware training, the 64-accelerator deployment recommendation, the cache hit rate and all three price tiers are quoted from Moonshot’s own Kimi K3 tech blog, which also hosts the demo renders and videos. The 1.4 TB and 5.6 TB footprints are our arithmetic on that parameter count, matched by TECHi of July 24, which is also the source for the eighteen-accelerator floor, the roughly 50 billion active parameter equivalent and the point that the license text does not yet exist. The MXFP4 hardware support note and the license-pending status are corroborated by a Hugging Face community write-up. The modified MIT precedent and its attribution threshold come from the earlier K2.7 Code release. Leaderboard placements are as reported by third parties and are not consistent between accounts.

Frequently asked questions

When are the Kimi K3 weights released?

Moonshot states it in its own K3 tech blog: the full model weights will be released by July 27, 2026. The hosted model has been live in the Kimi app and API since the evening of July 16. So the API came first and the download follows about eleven days later, which is the same order Moonshot used for earlier Kimi releases.

How much disk and memory do the Kimi K3 weights need?

About 1.4 TB. Moonshot published the parameter count, 2.8 trillion, and the format, MXFP4, which is half a byte per parameter, so 2.8 trillion times 0.5 bytes lands at roughly 1.4 TB. At 16-bit it would be around 5.6 TB. Moonshot itself does not state a file size anywhere, so treat 1.4 TB as arithmetic rather than a published spec. Weights are also only part of the bill: the KV cache for a 1M token context is extra.

What hardware does it take to actually serve Kimi K3?

More than most people assume. Just holding 1.4 TB of weights takes eighteen accelerators of 80 GB, and Moonshot recommends deploying on supernode configurations with 64 or more accelerators. MXFP4 is natively supported on NVIDIA Blackwell and AMD MI400 class silicon, which is another way of saying older cards will not give you the format for free. Renting inference is the realistic path for almost everybody.

What license do the Kimi K3 weights ship under?

Unpublished as of July 26. Coverage expects a modified MIT license, because that is what earlier Kimi models used, and the K2.7 Code variant carried a single attribution clause requiring products above roughly 100 million monthly active users or 20 million dollars in monthly revenue to display the model name. Expected is not published though. Until the license file lands next to the weights, nobody can tell you cleanly whether K3 is usable in your commercial product or what limits apply to training on its output.

Is the hosted Kimi K3 API cheaper than Claude Opus 5?

On list price, yes. Moonshot posts $0.30 per million tokens for cache-hit input, $3.00 for cache-miss input and $15.00 for output. Anthropic prices Opus 5 at $5 input and $25 output per million in standard mode. The cache tier is the part worth engineering around: Moonshot reports cache hit rates above 90 percent in coding workloads, and a ten times gap between hit and miss means your prompt prefix stability shows up on the invoice.

Tags: aikimillmmoonshotnewsopen-sourceself-hosting
Share197Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
Answer card: Apple released iOS 26.6 and iPadOS 26.6 on 27 July 2026 with a release note covering bug fixes, security updates and an optimized Spotlight index to prepare for iOS 27, the index the rebuilt Siri reads for personal context.

iOS 26.6 is out: the Spotlight index it quietly builds

27 July 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.