Somebody on the team says the new model is open source, so let's host it ourselves. Then somebody opens the licence file. Sometimes it's MIT. Sometimes it bars four regions outright, or wants a separate deal once your hosting business clears a revenue line. This page tracks 35 entries from 2026 and the part the headlines skip: are the weights downloadable, promised or API only, under which licence, at what size, and on what hardware. It's for sysadmins and devs who mean to run the thing themselves.
Last verified: 7 October 2026. On that day we read the repo record, commit history, licence file and file list on Hugging Face for all 26 entries with published weights, and the lab's own page for the other 9.
How to read this page
Every row carries a status pill. Confirmed means the publisher's own page says it and we read that page on 7 October. Reported means we only have press or a third-party index. Announced, not delivered means a lab has promised something that isn't in a repo yet. Rumour means a claim circulates that no lab source backs.
The Weights column is the one we care about. Published means a public, ungated repo with files and a licence file. Promised means a statement with no repo. Not released covers API-only and sales-only models. The Gap line under a status is where the headline and the repo disagree, and it happens more often than you'd hope. Sizes are decimal GB from the repo file list. The "what it needs to run" cell quotes the model card where it speaks, and says so when a number is our arithmetic or a community estimate. Benchmarks are left out on purpose, because they're lab-reported and they move weekly. The licence file doesn't.
The tracker
Three tables, ordered by what you're allowed to do, newest release first inside each. Read the licence column before the size column. Of the 26 entries with published weights, 19 sit under Apache 2.0, MIT or OpenMDW-1.1, and 7 don't.
Weights published, permissive licence
| Model | Lab | Parameters (total / active) | Licence | Weights | Size on disk or memory needed | What it needs to run | Released | Status | Source | Our write-up |
|---|---|---|---|---|---|---|---|---|---|---|
| Kolibri-1 | Aleph Alpha | 78B / 3.46B | Apache 2.0 | Published ungated, FP8 repo plus a BF16 repo | 78.9 GB (FP8 repo) | Card minimum: 2 A100 80 GB, 2 H100, 1 H200, 1 B200 or 1 B300 | 3 Oct 2026 | Confirmed Gap: 1M tokens is validated by the lab, but the native context is 262,144 | Model card Lab post | Kolibri-1 context |
| MiMo-V2.6 Pro and Flash (MOPD) | Xiaomi | Pro 1.02T / 42B; Flash 309B / 15B | MIT | Published ungated; RL build 21 Sep, MOPD retrain 27 Sep | Pro 574 GB, Flash 178 GB (repos) | Pro: vLLM with tensor parallel 8 (card), datacentre class | 21 Sep 2026 | Confirmed Gap: the API models were swapped to the retrained build on 25 Sep, same names, no way to pin the old one documented | Pro card Xiaomi post | MiMo retrain |
| Ternary Bonsai 2 27B | PrismML | 27.4B dense (Qwen3.8-27B base) | Apache 2.0 | Published ungated GGUF and MLX packs | 5.95 GB or 7.21 GB, plus 0.63 GB vision tower | 16 GB laptop or one 24 GB GPU, but only with PrismML's llama.cpp fork (stock builds reject or misread the files) | 17 Sep 2026 | Confirmed Gap: "98.2% kept" is a 14-benchmark average; agent rows keep about 75% | Model card Lab post | Bonsai 2 27B |
| Atria Dawn Preview | Shanghai AI Laboratory | 744B (753B in repo) / not stated | MIT | Published ungated, BF16 and FP8 | 1,507 GB BF16; 756 GB FP8 | SGLang 0.5.13.post1+ or vLLM 0.23.0+, GLM-5.2 recipes; FP8 is about ten 80 GB cards for weights alone | 11 Sep 2026 | Confirmed Gap: text only, the hosted endpoint rejects images | Model card Paper | Atria Dawn Preview |
| DeepSeek-V4.1-Flash | DeepSeek | 552B backbone (763B in repo) / 8B prefill, 16B decode | MIT | Published ungated | 510 GB, mostly 8-bit | Multi-GPU server. The repo ships a reference implementation (weights converted to tensor parallel 8), not a production engine; the card gives no vLLM or SGLang command | 10 Sep 2026 | Confirmed Gap: the V4 Pro retirement notice was reversed a day later; old deepseek-v4-flash names now route here | Model card DeepSeek post | V4.1-Flash and the Pro reversal |
| Hy4 preview | Tencent | 770B / 49B | Apache 2.0 | Published ungated, BF16 and FP8 | 814 GB FP8; 1,560 GB BF16 | Card serves FP8 at tensor parallel 8, about 102 GB of weights per card by our division, so eight 80 GB cards won't hold it | 28 Aug 2026 | Confirmed Gap: still a preview label, no final release | Model card Tencent post | Hy4 preview |
| GLM-5.3-Flash | Z.ai | 320B / 18B | MIT | Published ungated, FP8 and BF16 repos | 328 GB FP8; 643 GB BF16 | Recipes for SGLang, vLLM, TokenSpeed and KTransformers; a small cluster, not one card | 26 Aug 2026 | Confirmed Gap: it ran anonymously as Ox Alpha, with prompt retention, from 20 Aug | Model card | GLM-5.3-Flash |
| Qwen3.8-27B | Alibaba | 27B dense | Apache 2.0 | Published ungated, BF16 and FP8 repos | 55.6 GB BF16; 27.8 GB FP8 | About 28 GB at FP8, so a 48 GB card leaves room for the KV cache; 14 to 17 GB at 4-bit (community estimate) | 14 Aug 2026 | Confirmed Gap: the same week's 2.4T flagship is text only and custom-licensed | Model card FP8 card | Qwen3.8-27B |
| DeepSeek-V4-Pro-0813 | DeepSeek | 1.6T / 49B | MIT | Published ungated, 13 Aug | 893 GB | Card example: vLLM on one node of four GB300s | 13 Aug 2026 | Confirmed Gap: the API served this build on 12 Aug while only the April preview weights were public | Model card Change log | V4 Pro 0813 |
| Nemotron 3.5 Lightning | NVIDIA | 30B / 3B | OpenMDW-1.1 | Published ungated, NVFP4 and BF16, plus the RL dataset | 21.6 GB NVFP4; 65.8 GB BF16 | Card names one DGX Spark or one H100; RTX 5090 is listed with no memory floor or recipe | 11 Aug 2026 | Confirmed Gap: the card says up to 1M context; OpenRouter's paid endpoint lists 262K, its free one 1M | Model card NVIDIA post | Nemotron 3.5 Lightning |
| Muse Glimmer | Meta | 30B dense / 30B | Apache 2.0 | Published ungated, BF16, GGUF, ExecuTorch | 59.6 GB BF16; under 20 GB quantised | Card targets: 24 GB for the K-Quant-17GB build, 32 GB for K-Quant-Dynamic, 64 GB at full precision | 10 Aug 2026 | Confirmed Gap: the 20,000 tokens a second figure is NVIDIA's datacentre aggregate; Meta's own RTX 5090 figure is 233 with speculative decoding | Model card Meta post | Muse Glimmer |
| Shieldstral 1.0 3B | Mistral | 3B (3.85B with vision encoder) / dense | Apache 2.0 | Published ungated | 15.4 GB BF16 | One 16 GB GPU; vLLM, llama.cpp or Transformers | 4 Aug 2026 | Confirmed Gap: the HF card says 3B, the API card 3.8B; the repo's 3.85B includes the Pixtral vision encoder | Model card Mistral post | Shieldstral 1.0 |
| DeepSeek-V4-Flash-0731 | DeepSeek | 284B (304B in repo) / 13B | MIT | Published ungated; superseded | 167 GB | Card example: vLLM on one node of four GB300s | 31 Jul 2026 | Confirmed Gap: the API moved to V4.1-Flash on 10 Sep; this repo is the only way to keep the 0731 build | Model card | V4-Flash-0731 |
| Inkling-Small | Thinking Machines | 276B / 12B | Apache 2.0, plus a Model Acceptable Use Policy | Published ungated, BF16 and NVFP4 | 532 GB BF16; 171 GB NVFP4 | BF16: 600 GB of VRAM (four B300 or eight H200). NVFP4: 180 GB (one B300 or two H200) | 30 Jul 2026 | Confirmed Gap: a separate acceptable use policy sits beside the licence | Model card HF repo | Inkling-Small |
| Laguna S 2.1 | Poolside | 118B / 8B | OpenMDW-1.1 | Published ungated, BF16, FP8, INT4, NVFP4 | 235 GB BF16; about 59 GB at 4-bit (our arithmetic) | One DGX Spark (128 GB) at 4-bit; Poolside doesn't size the 1M-token cache | 21 Jul 2026 | Confirmed Gap: coverage said it beats models ten times its size; Poolside's own post says "in its weight class" | Model card Poolside post | Laguna S 2.1 |
| Inkling | Thinking Machines | 975B / 41B | Apache 2.0, plus a Model Acceptable Use Policy | Published ungated, BF16 and NVFP4 | 1,905 GB BF16; 592 GB NVFP4 | BF16: 2 TB of VRAM (eight B300 or sixteen H200). NVFP4: 600 GB (four B300 or eight H200) | 15 Jul 2026 | Confirmed Gap: a separate acceptable use policy sits beside the licence | Model card HF repo | Inkling |
| Hy3 | Tencent | 295B / 21B | Apache 2.0 | Published ungated, BF16 and FP8 | 598 GB BF16 | Eight GPUs of H20-3e class or larger memory (card); vLLM or SGLang | 6 Jul 2026 | Confirmed | Model card Licence | Hy3 |
| GLM-5.2 | Z.ai | 753B / not stated | MIT | Published ungated, BF16 and FP8 | 1,507 GB BF16 | Not stated on the card; SGLang and vLLM recipes. The base of Atria Dawn and Palmyra X6 | 16 Jun 2026 | Confirmed Gap: its successor GLM-5.3 is not MIT | Model card | GLM-5.2 |
| Qwen3.6-27B | Alibaba | 27B dense | Apache 2.0 | Published ungated, BF16 and FP8 | 55.6 GB BF16 | A 16 GB GPU at 4-bit through Ollama (our June write-up) | Apr 2026 | Confirmed Gap: this is the open Qwen next to the API-only 3.7 | Model card | Running Qwen locally |
Weights published, licence with strings attached
| Model | Lab | Parameters (total / active) | Licence | Weights | Size on disk or memory needed | What it needs to run | Released | Status | Source | Our write-up |
|---|---|---|---|---|---|---|---|---|---|---|
| FLUX 3 Action base | Black Forest Labs | 7B / not stated | FLUX Kommunity License v1.0: non-commercial; commercial use of outputs only under $5M annual revenue; production use needs a commercial licence from BFL | Published ungated; the text encoder is Apache 2.0 Qwen3-VL 4B | 25.4 GB | Card gives no memory floor; BFL's post discusses B200, H200, RTX 6000 Pro and RTX 5090 | 22 Sep 2026 | Confirmed Gap: July's "open weights later" meant FLUX 3 Dev, which has no date; the first FLUX 3 weights are a robot-action model | Model card Licence | No current write-up |
| Qwen-Image-2.1 | Alibaba | 7B DiT plus an 8B Qwen3-VL text encoder | Qwen Research License: research and evaluation only; commercial use needs a separate licence by email | Published ungated | 33.1 GB BF16; about 14 GB with INT8 or W4A8 repacks | 34 GB peak in BF16 at 1024 px (vLLM-Omni recipe); a 24 GB card with CPU offload | 20 Sep 2026 | Confirmed Gap: every earlier Qwen-Image checkpoint with weights was Apache 2.0 | Model card Licence | Qwen-Image-2.1 |
| GLM-5.3 | Z.ai | 753B / not stated | GLM-5.3 License: permissive, but Model-as-a-Service operators above $10B revenue (12 months) must pass a Z.ai security review before commercial use | Published ungated, FP8 and BF16 repos | 756 GB FP8; 1,507 GB BF16 | SGLang, vLLM, TokenSpeed, Transformers, KTransformers; Ascend NPU supported (card) | 27 Aug 2026 | Confirmed Gap: GLM-5.2 and 5.3-Flash are MIT, the flagship isn't; announced 14 Aug with "about two weeks" | Model card Licence | No current write-up |
| Qwen3.8-Flash-Next | Alibaba | 125B (about 180B stored) / 6B | Qwen Community License 1.0: any Model-as-a-Service or AI Work Assistant business needs a separate licence; name shown above 100M users or $20M monthly revenue | Published ungated, BF16 and FP8 | 186 GB FP8; 360 GB BF16 | A node, not a card. A third-party engine, Strata, runs 2-bit and 3-bit packs on 12 GB of VRAM plus 32 GB of RAM | 26 Aug 2026 | Confirmed Gap: some launch coverage said Apache 2.0; the repo says license "other" | Model card Licence | Qwen3.8-Flash-Next |
| Qwen3.8-Max (2.4T-A95B weights) | Alibaba | 2.4T / 95B | Qwen3.8-Max License: hosting or coding-assistant businesses above $50M revenue (12 months) need a separate licence; name shown above 100M users or $20M monthly revenue | Published ungated; text only, thinking can't be switched off | 4,892 GB BF16; 2,496 GB FP8 | Server class; vLLM, SGLang or TokenSpeed (card). Context 262,144 native, 1,010,000 extended | 12 Aug 2026 | Confirmed Gap: 4 Aug headlines said "open source"; the weights came 8 days later under a custom licence, without the vision and non-thinking features of hosted Qwen3.8-Max | Model card Licence | No current write-up |
| MiniMax H3 | MiniMax | 33B dense / 33B | MiniMax H3 Community License: no use of the weights or any output in the EU, UK, South Korea or US; written authorisation above $20M yearly revenue; "MiniMax H3" shown in the UI; no improving other models | Published ungated | 498 GB repo (all checkpoints); 42.5 GB smallest ComfyUI set | ComfyUI says an RTX 3060 class card completes a clip with offload; no timing published | 3 Aug 2026 | Confirmed Gap: the banner says "Open-Weights"; the licence excludes four regions, and the hosted API is exempt | Model card Licence | MiniMax H3 |
| Kimi K3 | Moonshot AI | 2.8T / 104B | Kimi K3 License: Model-as-a-Service above $20M revenue (12 months) needs a separate deal; name shown above 100M users or $20M monthly revenue; internal use exempt | Published ungated, MXFP4 | 1,561 GB | Twenty 80 GB cards hold the files alone; Moonshot recommends supernodes of 64 or more accelerators | 27 Jul 2026 | Confirmed Gap: coverage expected a "modified MIT"; the licence is its own text, and the card says 104B active | Model card Licence Moonshot post | No current write-up |
Four of these seven (Kimi K3, Qwen3.8-Max, GLM-5.3 and Qwen3.8-Flash-Next) are permissive text with a hosting clause bolted on. Using the weights inside your own company is fine under all four. Selling inference to third parties is where the clauses bite, and the trigger is your whole group's revenue, not what the hosting earns: a separate deal above $20 million for Kimi K3 and above $50 million for Qwen3.8-Max. Qwen3.8-Flash-Next sets no threshold: any such business needs a licence. At $10 billion GLM-5.3 asks for a security review rather than a deal. MiniMax H3 is the odd one out: it bars the weights and every output in four regions. One example isn't a trend, and I might be wrong, but anyone with customers there should treat H3 as an API product.
Promised, withheld or hosted only
| Model | Lab | Parameters (total / active) | Licence | Weights | Size on disk or memory needed | What it needs to run | Released | Status | Source | Our write-up |
|---|---|---|---|---|---|---|---|---|---|---|
| Reflection Beam | Reflection AI | 501B / 23B | Apache 2.0, promised | Promised "later this month", per the lab; no Hugging Face repo yet | Not published; every expert must still sit in memory | Waitlist only; no hardware guidance or quantised sizes | Announced 5 Oct 2026 | Announced, not delivered Gap: Apache 2.0 on the poster, a waitlist on the site | Lab post | Reflection Beam |
| Qwen3.8-Omni-Flash | Alibaba | Not published | None, hosted only | Not released API on Model Studio; no repo in the Qwen organisation | Not applicable | Hosted API; the last Omni weights are Qwen3-Omni-30B-A3B from Sep 2025 | 17 Sep 2026 | Confirmed: no weights Gap: the launch post links open-source companions, but not the model; the harness repo that returned 404 on 18 Sep loads now | Lab post Qwen on Hugging Face | Qwen3.8-Omni-Flash |
| Wan 3.0 | Alibaba | Not published | None, hosted only | Not released API only (wan3.0-video); no Wan 3.0 repo in the Wan-AI organisation | Not applicable | Model Studio API; endpoint, model and key must share a region | 24 Aug 2026 | Rumour: Apache 2.0 weights Gap: a claimed Apache 2.0 release in 1.3B and 14B sizes appears in no Alibaba source | API reference Wan-AI on Hugging Face | Wan 3.0 |
| Palmyra X6 | Writer | Not published | Proprietary, post-trained on MIT-licensed GLM-5.2 | Not released Writer customers only | Not applicable | Writer platform; no self-hosting path | 13 Aug 2026 | Confirmed: no weights Gap: the open base, GLM-5.2, is downloadable and the product built on it isn't | Writer post | Palmyra X6 |
| Genesis-Science-1 | US DOE with Arcee AI | Not published | Not stated | Promised weights, report and demos, no date | Not published | Nothing to run; contribution windows closed 14 and 25 Aug, the next roughly every three months | Announced 7 Aug 2026 | Announced, not delivered Gap: "DOE open-weight model" headlines described a portal and an application form | Arcee post DOE post | Genesis-Science-1 |
| FLUX 3 Video, Image and Dev | Black Forest Labs | Not published | Not stated for Dev | Promised Dev "later this year", no date; Video is early access | Not published | Early-access hosted Video; nothing to download except the Action model above | Announced 23 Jul 2026 | Announced, not delivered Gap: the open backbone is the part the FLUX name was built on, and it's the part still missing | BFL post BFL on Hugging Face | No current write-up |
| Qwen-Image 2.0 and 3.0 | Alibaba | Not published | None, hosted only | Not released 3.0 in Qwen Chat; no 2.0 or 3.0 repo exists | Not applicable | Hosted in Qwen Chat; use Qwen-Image-2.1 or the Apache 2.0 2512 build to self-host | Feb 2026 (2.0), 21 Jul 2026 (3.0) | Confirmed: no weights Gap: 3.0 also shipped without a model card or any benchmark score | Lab post Qwen on Hugging Face | Qwen-Image-2.1 and its siblings |
| Robostral Navigate | Mistral | 8B / not stated | None, sales-gated | Not released no weights, no public API; contact sales | Not applicable | Single RGB camera robot navigation; trained in simulation | 7 Jul 2026 | Confirmed: no weights Gap: a lab known for open weights shipped this one behind a sales form | Mistral post | Robostral Navigate |
| Qwen 3.7 Max and Plus | Alibaba | Not published | Proprietary | Not released API only; no Qwen3.7 repo in the Qwen organisation | Not applicable | Hosted on Model Studio; Qwen3.6-27B and Qwen3.8-27B are the open options | May and Jun 2026 | Reported: API only Gap: Alibaba's "API first, weights weeks later" habit did not repeat for 3.7 | The Batch Qwen on Hugging Face | Qwen 3.7 locally |
What changed since
Newest first, limited to changes in weight availability or licence terms.
- 5 October 2026. Reflection AI announces Beam, 501B total and 23B active, and promises Apache 2.0 weights later in October. Today it's a waitlist and there's no Hugging Face repo. Reflection, our write-up.
- 3 October 2026. Aleph Alpha publishes Kolibri-1 under Apache 2.0, FP8 and BF16 repos. Model card.
- 27 September 2026. Xiaomi publishes retrained MOPD checkpoints of MiMo-V2.6, still MIT. The API models were swapped two days earlier under the same names. Model card.
- 22 to 23 September 2026. Black Forest Labs uploads FLUX 3 Action base under the FLUX Kommunity License: the first FLUX 3 weights, non-commercial, with a $5M revenue carve-out for outputs. FLUX 3 Dev still has no date. Licence.
- 20 September 2026. Qwen-Image-2.1 weights land under the Qwen Research License, research and evaluation only. Every earlier Qwen-Image checkpoint with weights was Apache 2.0. Licence.
- 17 September 2026. PrismML ships Ternary Bonsai 2 27B, Apache 2.0, from 5.95 GB, readable only by its own llama.cpp fork. Model card.
- 11 September 2026. The Shanghai AI Laboratory publishes Atria Dawn Preview under MIT, on the GLM-5.2 base. Text only. Model card.
- 10 to 11 September 2026. DeepSeek releases V4.1-Flash under MIT, announces that V4 Pro will be routed into it, then keeps Pro on the API a day later. Vision-Exp, posted to Hugging Face on 31 August, is folded in. DeepSeek change log, our write-up.
- 27 August 2026. Z.ai publishes GLM-5.3 weights 13 days after its 14 August launch announcement (its docs changelog dates the API entry 18 August), under a new GLM-5.3 License instead of MIT. Licence.
- 26 August 2026. Z.ai opens GLM-5.3-Flash under MIT and the anonymous Ox Alpha listing goes away. Alibaba opens Qwen3.8-Flash-Next under the Qwen Community License 1.0, which some launch coverage called Apache 2.0. GLM-5.3-Flash, Flash-Next.
- 24 August 2026. Alibaba rolls out Wan 3.0 as an API only. The Apache 2.0 weights some posts described appear in no Alibaba source. Our write-up.
- 14 August 2026. Qwen3.8-27B appears under Apache 2.0, with vision. Model card, our write-up.
- 13 August 2026. DeepSeek publishes V4-Pro-0813 under MIT, a day after the API switched to that build. Model card.
- 12 August 2026. Qwen3.8-2.4T-A95B gets its LICENSE file, the Qwen3.8-Max License. Alibaba had promised weights for the week after its 3 August announcement. Licence.
- 3 August 2026. MiniMax puts H3 on Hugging Face under a licence that bars use of the weights and of every output in the EU, UK, South Korea and US. Licence.
- 27 July 2026. Moonshot publishes Kimi K3 weights on its own deadline, 11 days after the API, under the Kimi K3 License rather than the modified MIT that coverage expected. Licence.
- 23 July 2026. Black Forest Labs announces FLUX 3 and promises the open FLUX 3 Dev backbone "later this year". BFL.
Frequently asked questions
Is "open weights" the same as open source on this page?
No. Of the 26 entries with published weights, 19 sit under Apache 2.0, MIT or OpenMDW-1.1 and 7 don't. Those 7 add a hosting clause, a non-commercial limit or a territorial ban. Even the Apache 2.0 ones can have a separate acceptable use policy next to them, as both Inkling models do.
What can I run on a single GPU?
Shieldstral (16 GB), Nemotron 3.5 Lightning in NVFP4 (21.6 GB; the card names one H100 or one DGX Spark), Muse Glimmer and Ternary Bonsai 2 27B (24 GB cards), Qwen3.8-27B (about 28 GB at FP8) and Laguna S 2.1 at 4-bit on a DGX Spark. Inkling-Small needs a B300 even at NVFP4. Everything else here is a node, not a card.
Which promised weights should I plan around?
Only the dated ones. Every dated promise in the log landed. Kimi K3 hit its 27 July deadline and GLM-5.3 arrived 13 days after its 14 August launch, inside the "about two weeks" Z.ai had named. Qwen3.8-Max came inside the week Alibaba named. The undated ones are still open: FLUX 3 Dev since 23 July, Genesis-Science-1 since August. Reflection's "later this month" is the next test.
Should I mirror weights I depend on?
If a build needs them, yes. Pin the revision hash and keep a copy you control. DeepSeek swapped the build behind the same API name twice this summer, and V4-Flash-0731 now survives only as a Hugging Face repo. We covered the policy angle in our note on the Open Weights letter.
How do you check these rows?
For published weights we read the Hugging Face record, commit history, LICENSE file and file list. For the rest we read the lab's page and search its Hugging Face organisation. When a row disagrees with one of our older articles, the repo wins and the article isn't linked until it's fixed.
Sources
Hugging Face model API and commit history for each repo, read on 7 October 2026, for example Kimi K3, GLM-5.3, Qwen3.8-2.4T-A95B and MiniMax H3. Lab pages: Thinking Machines, Moonshot, Reflection AI, Black Forest Labs, Arcee AI and the DeepSeek change log. Third party, only where a lab page is silent: The Batch on Qwen3.7.