Open-weight models: licence, size and hardware, row by row

Somebody on the team says the new model is open source, so let's host it ourselves. Then somebody opens the licence file. Sometimes it's MIT. Sometimes it bars four regions outright, or wants a separate deal once your hosting business clears a revenue line. This page tracks 35 entries from 2026 and the part the headlines skip: are the weights downloadable, promised or API only, under which licence, at what size, and on what hardware. It's for sysadmins and devs who mean to run the thing themselves.

Last verified: 7 October 2026. On that day we read the repo record, commit history, licence file and file list on Hugging Face for all 26 entries with published weights, and the lab's own page for the other 9.

How to read this page

Every row carries a status pill. Confirmed means the publisher's own page says it and we read that page on 7 October. Reported means we only have press or a third-party index. Announced, not delivered means a lab has promised something that isn't in a repo yet. Rumour means a claim circulates that no lab source backs.

The Weights column is the one we care about. Published means a public, ungated repo with files and a licence file. Promised means a statement with no repo. Not released covers API-only and sales-only models. The Gap line under a status is where the headline and the repo disagree, and it happens more often than you'd hope. Sizes are decimal GB from the repo file list. The "what it needs to run" cell quotes the model card where it speaks, and says so when a number is our arithmetic or a community estimate. Benchmarks are left out on purpose, because they're lab-reported and they move weekly. The licence file doesn't.

The tracker

Three tables, ordered by what you're allowed to do, newest release first inside each. Read the licence column before the size column. Of the 26 entries with published weights, 19 sit under Apache 2.0, MIT or OpenMDW-1.1, and 7 don't.

Weights published, permissive licence

Weights you can download today under Apache 2.0, MIT or OpenMDW-1.1. Newest release first. Sizes are decimal GB summed from the repo file list on 7 October 2026 unless the cell says otherwise.
ModelLabParameters (total / active)LicenceWeightsSize on disk or memory neededWhat it needs to runReleasedStatusSourceOur write-up
Kolibri-1Aleph Alpha78B / 3.46BApache 2.0Published
ungated, FP8 repo plus a BF16 repo
78.9 GB (FP8 repo)Card minimum: 2 A100 80 GB, 2 H100, 1 H200, 1 B200 or 1 B3003 Oct 2026Confirmed
Gap: 1M tokens is validated by the lab, but the native context is 262,144
Model card
Lab post
Kolibri-1 context
MiMo-V2.6 Pro and Flash (MOPD)XiaomiPro 1.02T / 42B; Flash 309B / 15BMITPublished
ungated; RL build 21 Sep, MOPD retrain 27 Sep
Pro 574 GB, Flash 178 GB (repos)Pro: vLLM with tensor parallel 8 (card), datacentre class21 Sep 2026Confirmed
Gap: the API models were swapped to the retrained build on 25 Sep, same names, no way to pin the old one documented
Pro card
Xiaomi post
MiMo retrain
Ternary Bonsai 2 27BPrismML27.4B dense (Qwen3.8-27B base)Apache 2.0Published
ungated GGUF and MLX packs
5.95 GB or 7.21 GB, plus 0.63 GB vision tower16 GB laptop or one 24 GB GPU, but only with PrismML's llama.cpp fork (stock builds reject or misread the files)17 Sep 2026Confirmed
Gap: "98.2% kept" is a 14-benchmark average; agent rows keep about 75%
Model card
Lab post
Bonsai 2 27B
Atria Dawn PreviewShanghai AI Laboratory744B (753B in repo) / not statedMITPublished
ungated, BF16 and FP8
1,507 GB BF16; 756 GB FP8SGLang 0.5.13.post1+ or vLLM 0.23.0+, GLM-5.2 recipes; FP8 is about ten 80 GB cards for weights alone11 Sep 2026Confirmed
Gap: text only, the hosted endpoint rejects images
Model card
Paper
Atria Dawn Preview
DeepSeek-V4.1-FlashDeepSeek552B backbone (763B in repo) / 8B prefill, 16B decodeMITPublished
ungated
510 GB, mostly 8-bitMulti-GPU server. The repo ships a reference implementation (weights converted to tensor parallel 8), not a production engine; the card gives no vLLM or SGLang command10 Sep 2026Confirmed
Gap: the V4 Pro retirement notice was reversed a day later; old deepseek-v4-flash names now route here
Model card
DeepSeek post
V4.1-Flash and the Pro reversal
Hy4 previewTencent770B / 49BApache 2.0Published
ungated, BF16 and FP8
814 GB FP8; 1,560 GB BF16Card serves FP8 at tensor parallel 8, about 102 GB of weights per card by our division, so eight 80 GB cards won't hold it28 Aug 2026Confirmed
Gap: still a preview label, no final release
Model card
Tencent post
Hy4 preview
GLM-5.3-FlashZ.ai320B / 18BMITPublished
ungated, FP8 and BF16 repos
328 GB FP8; 643 GB BF16Recipes for SGLang, vLLM, TokenSpeed and KTransformers; a small cluster, not one card26 Aug 2026Confirmed
Gap: it ran anonymously as Ox Alpha, with prompt retention, from 20 Aug
Model cardGLM-5.3-Flash
Qwen3.8-27BAlibaba27B denseApache 2.0Published
ungated, BF16 and FP8 repos
55.6 GB BF16; 27.8 GB FP8About 28 GB at FP8, so a 48 GB card leaves room for the KV cache; 14 to 17 GB at 4-bit (community estimate)14 Aug 2026Confirmed
Gap: the same week's 2.4T flagship is text only and custom-licensed
Model card
FP8 card
Qwen3.8-27B
DeepSeek-V4-Pro-0813DeepSeek1.6T / 49BMITPublished
ungated, 13 Aug
893 GBCard example: vLLM on one node of four GB300s13 Aug 2026Confirmed
Gap: the API served this build on 12 Aug while only the April preview weights were public
Model card
Change log
V4 Pro 0813
Nemotron 3.5 LightningNVIDIA30B / 3BOpenMDW-1.1Published
ungated, NVFP4 and BF16, plus the RL dataset
21.6 GB NVFP4; 65.8 GB BF16Card names one DGX Spark or one H100; RTX 5090 is listed with no memory floor or recipe11 Aug 2026Confirmed
Gap: the card says up to 1M context; OpenRouter's paid endpoint lists 262K, its free one 1M
Model card
NVIDIA post
Nemotron 3.5 Lightning
Muse GlimmerMeta30B dense / 30BApache 2.0Published
ungated, BF16, GGUF, ExecuTorch
59.6 GB BF16; under 20 GB quantisedCard targets: 24 GB for the K-Quant-17GB build, 32 GB for K-Quant-Dynamic, 64 GB at full precision10 Aug 2026Confirmed
Gap: the 20,000 tokens a second figure is NVIDIA's datacentre aggregate; Meta's own RTX 5090 figure is 233 with speculative decoding
Model card
Meta post
Muse Glimmer
Shieldstral 1.0 3BMistral3B (3.85B with vision encoder) / denseApache 2.0Published
ungated
15.4 GB BF16One 16 GB GPU; vLLM, llama.cpp or Transformers4 Aug 2026Confirmed
Gap: the HF card says 3B, the API card 3.8B; the repo's 3.85B includes the Pixtral vision encoder
Model card
Mistral post
Shieldstral 1.0
DeepSeek-V4-Flash-0731DeepSeek284B (304B in repo) / 13BMITPublished
ungated; superseded
167 GBCard example: vLLM on one node of four GB300s31 Jul 2026Confirmed
Gap: the API moved to V4.1-Flash on 10 Sep; this repo is the only way to keep the 0731 build
Model cardV4-Flash-0731
Inkling-SmallThinking Machines276B / 12BApache 2.0, plus a Model Acceptable Use PolicyPublished
ungated, BF16 and NVFP4
532 GB BF16; 171 GB NVFP4BF16: 600 GB of VRAM (four B300 or eight H200). NVFP4: 180 GB (one B300 or two H200)30 Jul 2026Confirmed
Gap: a separate acceptable use policy sits beside the licence
Model card
HF repo
Inkling-Small
Laguna S 2.1Poolside118B / 8BOpenMDW-1.1Published
ungated, BF16, FP8, INT4, NVFP4
235 GB BF16; about 59 GB at 4-bit (our arithmetic)One DGX Spark (128 GB) at 4-bit; Poolside doesn't size the 1M-token cache21 Jul 2026Confirmed
Gap: coverage said it beats models ten times its size; Poolside's own post says "in its weight class"
Model card
Poolside post
Laguna S 2.1
InklingThinking Machines975B / 41BApache 2.0, plus a Model Acceptable Use PolicyPublished
ungated, BF16 and NVFP4
1,905 GB BF16; 592 GB NVFP4BF16: 2 TB of VRAM (eight B300 or sixteen H200). NVFP4: 600 GB (four B300 or eight H200)15 Jul 2026Confirmed
Gap: a separate acceptable use policy sits beside the licence
Model card
HF repo
Inkling
Hy3Tencent295B / 21BApache 2.0Published
ungated, BF16 and FP8
598 GB BF16Eight GPUs of H20-3e class or larger memory (card); vLLM or SGLang6 Jul 2026ConfirmedModel card
Licence
Hy3
GLM-5.2Z.ai753B / not statedMITPublished
ungated, BF16 and FP8
1,507 GB BF16Not stated on the card; SGLang and vLLM recipes. The base of Atria Dawn and Palmyra X616 Jun 2026Confirmed
Gap: its successor GLM-5.3 is not MIT
Model cardGLM-5.2
Qwen3.6-27BAlibaba27B denseApache 2.0Published
ungated, BF16 and FP8
55.6 GB BF16A 16 GB GPU at 4-bit through Ollama (our June write-up)Apr 2026Confirmed
Gap: this is the open Qwen next to the API-only 3.7
Model cardRunning Qwen locally

Weights published, licence with strings attached

Weights you can download today, but the licence adds a revenue trigger, a hosting clause, a non-commercial limit or a territorial ban. Quoted from the LICENSE file in each repo, read on 7 October 2026.
ModelLabParameters (total / active)LicenceWeightsSize on disk or memory neededWhat it needs to runReleasedStatusSourceOur write-up
FLUX 3 Action baseBlack Forest Labs7B / not statedFLUX Kommunity License v1.0: non-commercial; commercial use of outputs only under $5M annual revenue; production use needs a commercial licence from BFLPublished
ungated; the text encoder is Apache 2.0 Qwen3-VL 4B
25.4 GBCard gives no memory floor; BFL's post discusses B200, H200, RTX 6000 Pro and RTX 509022 Sep 2026Confirmed
Gap: July's "open weights later" meant FLUX 3 Dev, which has no date; the first FLUX 3 weights are a robot-action model
Model card
Licence
No current write-up
Qwen-Image-2.1Alibaba7B DiT plus an 8B Qwen3-VL text encoderQwen Research License: research and evaluation only; commercial use needs a separate licence by emailPublished
ungated
33.1 GB BF16; about 14 GB with INT8 or W4A8 repacks34 GB peak in BF16 at 1024 px (vLLM-Omni recipe); a 24 GB card with CPU offload20 Sep 2026Confirmed
Gap: every earlier Qwen-Image checkpoint with weights was Apache 2.0
Model card
Licence
Qwen-Image-2.1
GLM-5.3Z.ai753B / not statedGLM-5.3 License: permissive, but Model-as-a-Service operators above $10B revenue (12 months) must pass a Z.ai security review before commercial usePublished
ungated, FP8 and BF16 repos
756 GB FP8; 1,507 GB BF16SGLang, vLLM, TokenSpeed, Transformers, KTransformers; Ascend NPU supported (card)27 Aug 2026Confirmed
Gap: GLM-5.2 and 5.3-Flash are MIT, the flagship isn't; announced 14 Aug with "about two weeks"
Model card
Licence
No current write-up
Qwen3.8-Flash-NextAlibaba125B (about 180B stored) / 6BQwen Community License 1.0: any Model-as-a-Service or AI Work Assistant business needs a separate licence; name shown above 100M users or $20M monthly revenuePublished
ungated, BF16 and FP8
186 GB FP8; 360 GB BF16A node, not a card. A third-party engine, Strata, runs 2-bit and 3-bit packs on 12 GB of VRAM plus 32 GB of RAM26 Aug 2026Confirmed
Gap: some launch coverage said Apache 2.0; the repo says license "other"
Model card
Licence
Qwen3.8-Flash-Next
Qwen3.8-Max (2.4T-A95B weights)Alibaba2.4T / 95BQwen3.8-Max License: hosting or coding-assistant businesses above $50M revenue (12 months) need a separate licence; name shown above 100M users or $20M monthly revenuePublished
ungated; text only, thinking can't be switched off
4,892 GB BF16; 2,496 GB FP8Server class; vLLM, SGLang or TokenSpeed (card). Context 262,144 native, 1,010,000 extended12 Aug 2026Confirmed
Gap: 4 Aug headlines said "open source"; the weights came 8 days later under a custom licence, without the vision and non-thinking features of hosted Qwen3.8-Max
Model card
Licence
No current write-up
MiniMax H3MiniMax33B dense / 33BMiniMax H3 Community License: no use of the weights or any output in the EU, UK, South Korea or US; written authorisation above $20M yearly revenue; "MiniMax H3" shown in the UI; no improving other modelsPublished
ungated
498 GB repo (all checkpoints); 42.5 GB smallest ComfyUI setComfyUI says an RTX 3060 class card completes a clip with offload; no timing published3 Aug 2026Confirmed
Gap: the banner says "Open-Weights"; the licence excludes four regions, and the hosted API is exempt
Model card
Licence
MiniMax H3
Kimi K3Moonshot AI2.8T / 104BKimi K3 License: Model-as-a-Service above $20M revenue (12 months) needs a separate deal; name shown above 100M users or $20M monthly revenue; internal use exemptPublished
ungated, MXFP4
1,561 GBTwenty 80 GB cards hold the files alone; Moonshot recommends supernodes of 64 or more accelerators27 Jul 2026Confirmed
Gap: coverage expected a "modified MIT"; the licence is its own text, and the card says 104B active
Model card
Licence
Moonshot post
No current write-up

Four of these seven (Kimi K3, Qwen3.8-Max, GLM-5.3 and Qwen3.8-Flash-Next) are permissive text with a hosting clause bolted on. Using the weights inside your own company is fine under all four. Selling inference to third parties is where the clauses bite, and the trigger is your whole group's revenue, not what the hosting earns: a separate deal above $20 million for Kimi K3 and above $50 million for Qwen3.8-Max. Qwen3.8-Flash-Next sets no threshold: any such business needs a licence. At $10 billion GLM-5.3 asks for a security review rather than a deal. MiniMax H3 is the odd one out: it bars the weights and every output in four regions. One example isn't a trend, and I might be wrong, but anyone with customers there should treat H3 as an API product.

Promised, withheld or hosted only

No download today. Newest first. "Promised" is a lab statement without a repo; "Not released" covers API-only and sales-only models.
ModelLabParameters (total / active)LicenceWeightsSize on disk or memory neededWhat it needs to runReleasedStatusSourceOur write-up
Reflection BeamReflection AI501B / 23BApache 2.0, promisedPromised
"later this month", per the lab; no Hugging Face repo yet
Not published; every expert must still sit in memoryWaitlist only; no hardware guidance or quantised sizesAnnounced 5 Oct 2026Announced, not delivered
Gap: Apache 2.0 on the poster, a waitlist on the site
Lab postReflection Beam
Qwen3.8-Omni-FlashAlibabaNot publishedNone, hosted onlyNot released
API on Model Studio; no repo in the Qwen organisation
Not applicableHosted API; the last Omni weights are Qwen3-Omni-30B-A3B from Sep 202517 Sep 2026Confirmed: no weights
Gap: the launch post links open-source companions, but not the model; the harness repo that returned 404 on 18 Sep loads now
Lab post
Qwen on Hugging Face
Qwen3.8-Omni-Flash
Wan 3.0AlibabaNot publishedNone, hosted onlyNot released
API only (wan3.0-video); no Wan 3.0 repo in the Wan-AI organisation
Not applicableModel Studio API; endpoint, model and key must share a region24 Aug 2026Rumour: Apache 2.0 weights
Gap: a claimed Apache 2.0 release in 1.3B and 14B sizes appears in no Alibaba source
API reference
Wan-AI on Hugging Face
Wan 3.0
Palmyra X6WriterNot publishedProprietary, post-trained on MIT-licensed GLM-5.2Not released
Writer customers only
Not applicableWriter platform; no self-hosting path13 Aug 2026Confirmed: no weights
Gap: the open base, GLM-5.2, is downloadable and the product built on it isn't
Writer postPalmyra X6
Genesis-Science-1US DOE with Arcee AINot publishedNot statedPromised
weights, report and demos, no date
Not publishedNothing to run; contribution windows closed 14 and 25 Aug, the next roughly every three monthsAnnounced 7 Aug 2026Announced, not delivered
Gap: "DOE open-weight model" headlines described a portal and an application form
Arcee post
DOE post
Genesis-Science-1
FLUX 3 Video, Image and DevBlack Forest LabsNot publishedNot stated for DevPromised
Dev "later this year", no date; Video is early access
Not publishedEarly-access hosted Video; nothing to download except the Action model aboveAnnounced 23 Jul 2026Announced, not delivered
Gap: the open backbone is the part the FLUX name was built on, and it's the part still missing
BFL post
BFL on Hugging Face
No current write-up
Qwen-Image 2.0 and 3.0AlibabaNot publishedNone, hosted onlyNot released
3.0 in Qwen Chat; no 2.0 or 3.0 repo exists
Not applicableHosted in Qwen Chat; use Qwen-Image-2.1 or the Apache 2.0 2512 build to self-hostFeb 2026 (2.0), 21 Jul 2026 (3.0)Confirmed: no weights
Gap: 3.0 also shipped without a model card or any benchmark score
Lab post
Qwen on Hugging Face
Qwen-Image-2.1 and its siblings
Robostral NavigateMistral8B / not statedNone, sales-gatedNot released
no weights, no public API; contact sales
Not applicableSingle RGB camera robot navigation; trained in simulation7 Jul 2026Confirmed: no weights
Gap: a lab known for open weights shipped this one behind a sales form
Mistral postRobostral Navigate
Qwen 3.7 Max and PlusAlibabaNot publishedProprietaryNot released
API only; no Qwen3.7 repo in the Qwen organisation
Not applicableHosted on Model Studio; Qwen3.6-27B and Qwen3.8-27B are the open optionsMay and Jun 2026Reported: API only
Gap: Alibaba's "API first, weights weeks later" habit did not repeat for 3.7
The Batch
Qwen on Hugging Face
Qwen 3.7 locally

What changed since

Newest first, limited to changes in weight availability or licence terms.

  • 5 October 2026. Reflection AI announces Beam, 501B total and 23B active, and promises Apache 2.0 weights later in October. Today it's a waitlist and there's no Hugging Face repo. Reflection, our write-up.
  • 3 October 2026. Aleph Alpha publishes Kolibri-1 under Apache 2.0, FP8 and BF16 repos. Model card.
  • 27 September 2026. Xiaomi publishes retrained MOPD checkpoints of MiMo-V2.6, still MIT. The API models were swapped two days earlier under the same names. Model card.
  • 22 to 23 September 2026. Black Forest Labs uploads FLUX 3 Action base under the FLUX Kommunity License: the first FLUX 3 weights, non-commercial, with a $5M revenue carve-out for outputs. FLUX 3 Dev still has no date. Licence.
  • 20 September 2026. Qwen-Image-2.1 weights land under the Qwen Research License, research and evaluation only. Every earlier Qwen-Image checkpoint with weights was Apache 2.0. Licence.
  • 17 September 2026. PrismML ships Ternary Bonsai 2 27B, Apache 2.0, from 5.95 GB, readable only by its own llama.cpp fork. Model card.
  • 11 September 2026. The Shanghai AI Laboratory publishes Atria Dawn Preview under MIT, on the GLM-5.2 base. Text only. Model card.
  • 10 to 11 September 2026. DeepSeek releases V4.1-Flash under MIT, announces that V4 Pro will be routed into it, then keeps Pro on the API a day later. Vision-Exp, posted to Hugging Face on 31 August, is folded in. DeepSeek change log, our write-up.
  • 27 August 2026. Z.ai publishes GLM-5.3 weights 13 days after its 14 August launch announcement (its docs changelog dates the API entry 18 August), under a new GLM-5.3 License instead of MIT. Licence.
  • 26 August 2026. Z.ai opens GLM-5.3-Flash under MIT and the anonymous Ox Alpha listing goes away. Alibaba opens Qwen3.8-Flash-Next under the Qwen Community License 1.0, which some launch coverage called Apache 2.0. GLM-5.3-Flash, Flash-Next.
  • 24 August 2026. Alibaba rolls out Wan 3.0 as an API only. The Apache 2.0 weights some posts described appear in no Alibaba source. Our write-up.
  • 14 August 2026. Qwen3.8-27B appears under Apache 2.0, with vision. Model card, our write-up.
  • 13 August 2026. DeepSeek publishes V4-Pro-0813 under MIT, a day after the API switched to that build. Model card.
  • 12 August 2026. Qwen3.8-2.4T-A95B gets its LICENSE file, the Qwen3.8-Max License. Alibaba had promised weights for the week after its 3 August announcement. Licence.
  • 3 August 2026. MiniMax puts H3 on Hugging Face under a licence that bars use of the weights and of every output in the EU, UK, South Korea and US. Licence.
  • 27 July 2026. Moonshot publishes Kimi K3 weights on its own deadline, 11 days after the API, under the Kimi K3 License rather than the modified MIT that coverage expected. Licence.
  • 23 July 2026. Black Forest Labs announces FLUX 3 and promises the open FLUX 3 Dev backbone "later this year". BFL.

Frequently asked questions

Is "open weights" the same as open source on this page?

No. Of the 26 entries with published weights, 19 sit under Apache 2.0, MIT or OpenMDW-1.1 and 7 don't. Those 7 add a hosting clause, a non-commercial limit or a territorial ban. Even the Apache 2.0 ones can have a separate acceptable use policy next to them, as both Inkling models do.

What can I run on a single GPU?

Shieldstral (16 GB), Nemotron 3.5 Lightning in NVFP4 (21.6 GB; the card names one H100 or one DGX Spark), Muse Glimmer and Ternary Bonsai 2 27B (24 GB cards), Qwen3.8-27B (about 28 GB at FP8) and Laguna S 2.1 at 4-bit on a DGX Spark. Inkling-Small needs a B300 even at NVFP4. Everything else here is a node, not a card.

Which promised weights should I plan around?

Only the dated ones. Every dated promise in the log landed. Kimi K3 hit its 27 July deadline and GLM-5.3 arrived 13 days after its 14 August launch, inside the "about two weeks" Z.ai had named. Qwen3.8-Max came inside the week Alibaba named. The undated ones are still open: FLUX 3 Dev since 23 July, Genesis-Science-1 since August. Reflection's "later this month" is the next test.

Should I mirror weights I depend on?

If a build needs them, yes. Pin the revision hash and keep a copy you control. DeepSeek swapped the build behind the same API name twice this summer, and V4-Flash-0731 now survives only as a Hugging Face repo. We covered the policy angle in our note on the Open Weights letter.

How do you check these rows?

For published weights we read the Hugging Face record, commit history, LICENSE file and file list. For the rest we read the lab's page and search its Hugging Face organisation. When a row disagrees with one of our older articles, the repo wins and the article isn't linked until it's fixed.

Sources

Hugging Face model API and commit history for each repo, read on 7 October 2026, for example Kimi K3, GLM-5.3, Qwen3.8-2.4T-A95B and MiniMax H3. Lab pages: Thinking Machines, Moonshot, Reflection AI, Black Forest Labs, Arcee AI and the DeepSeek change log. Third party, only where a lab page is silent: The Batch on Qwen3.7.