• Latest
  • Trending
  • All
Thinking Machines Lab announcement artwork for Inkling-Small, hand-drawn black loops winding across a cream graph-paper grid, carrying no text or product imagery.

Inkling-Small out-codes the 975B flagship at 276B

3 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

7 September 2026
Answer card stating that Mullvad announced on 3 September 2026 that it is shutting down its public encrypted domain name system servers on 2 November 2026 and sponsoring the Quad9 Foundation instead, with 194.242.2.2 and its five sibling addresses all going away, and virtual private network customers unaffected.

Mullvad’s DNS servers go dark on 2 November, and Quad9 blocks no ads

5 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Friday, September 18, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Inkling-Small out-codes the 975B flagship at 276B

by stephane
3 September 2026
in Dev
0
Thinking Machines Lab announcement artwork for Inkling-Small, hand-drawn black loops winding across a cream graph-paper grid, carrying no text or product imagery.
494
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Two weeks ago we said the honest problem with Inkling was that you probably couldn't run it. Two terabytes of VRAM will do that. Thinking Machines seems to have agreed, because on 30 July it shipped Inkling-Small: 276B total parameters, 12B active, Apache 2.0, a million tokens of context. Then the part we didn't expect. It beats its own 975B parent on SWE-bench Verified and on Terminal-Bench 2.1, while asking for roughly a third of the memory. Free lunch? No. The same shrink that gained it two and a half points on SWE-bench cost it more than half its factual recall, and a third of its score on multi-turn banking tool use. Parameters store facts. Take 700B of them away and the facts leave too.

The short answer

Inkling-Small is a quarter the size of the Inkling flagship and better at agentic coding than it. The trade is recall: factual questions and long tool-use conversations both fall hard. If your workload retrieves what it needs, this is the better model and a third of the metal. If it expects the weights to know things, it isn’t.

276B12B active, was 41B
80.2SWE-bench Verified, beats the 975B
180 GBVRAM floor, quantized
Answer card: Thinking Machines released Inkling-Small on 30 July 2026, a 276B mixture-of-experts model with 12B active parameters under Apache 2.0, scoring 80.2 percent on SWE-bench Verified against 77.6 for the 975B Inkling, running on 180 GB of VRAM quantized instead of 600 GB, with SimpleQA Verified dropping from 43.9 to 20.6 percent.
A smaller model that codes better is not the usual direction of travel.

What actually shipped

Thinking Machines published Inkling-Small on 30 July, weights on Hugging Face, Apache 2.0, with a Model Acceptable Use Policy sitting alongside the licence. The architecture is a 42-layer decoder-only transformer with a sparse mixture-of-experts feed-forward stack. Each token goes to 6 of 256 experts, plus 2 shared experts that always fire. Attention alternates local and global layers.

276B total, 12B active. The 975B flagship from July ran 41B active across 66 layers, so this isn’t a distillation of the big one so much as a smaller sibling built on the same recipe.

Everything else carried over. A context window up to 1M tokens. Text, images and audio in, text out, with the same encoder-free multimodal handling. The thinking-effort dial that we liked in July is still there, still exposed as named levels, and it matters more here: at 12B active per token, the reasoning tokens you spend cost about a third of what they did.

Thinking Machines Lab announcement artwork for Inkling-Small, hand-drawn black loops on a cream graph-paper grid.

Image: Thinking Machines Lab

It gained code and lost facts

Here’s where it gets interesting. The lab’s own headline claim is that Inkling-Small reaches comparable performance at a quarter the size, and that it edges the flagship on Humanity’s Last Exam, 31.6% against 29.7%.

That undersells one half of the story and hides the other.

Diagram comparing Inkling-Small against the 975B Inkling on five benchmarks: SWE-bench Verified 80.2 against 77.6, Terminal-Bench 2.1 64.7 against 63.8, IFBench 82.2 against 79.8, SimpleQA Verified 20.6 against 43.9, and Tau 3 Banking 15.5 against 23.7.
Three wins and two collapses, from the same lab's two model cards.

The wins are all in the same family. Agentic coding, terminal work, instruction following. SWE-bench Verified goes 77.6 to 80.2, Terminal-Bench 2.1 goes 63.8 to 64.7, IFBench 79.8 to 82.2. None of those are huge, but they all point the same way, and a smaller model beating a bigger one from the same lab on any of them is worth a raised eyebrow.

Then SimpleQA Verified: 43.9 down to 20.6. That’s less than half. Tau 3 Banking, which measures multi-turn tool use inside a policy, goes 23.7 to 15.5.

We think the read is fairly plain. Skills that live in the reasoning loop survived the cut. Anything that needs the weights themselves to hold a fact did not, because that’s exactly what the missing 700B parameters were storing. Nice to see it this cleanly isolated, honestly, since both numbers came from the same lab in the same fortnight.

Which one you care about depends entirely on what sits in front of the model. Give it a codebase, a shell and a search tool, and 20.6 on SimpleQA barely registers. Ask it to answer from memory in a support flow and you’ve just halved your accuracy to save some GPUs.

Tau 3 is the number we’d actually worry about. Multi-turn tool use is what everyone is building right now, and 15.5% is low in absolute terms, not just relative to the flagship. Neither model looks good there.

Small is doing a lot of work in that name

Comparison chart of aggregated VRAM required by each checkpoint: Inkling BF16 at 2,000 GB, Inkling NVFP4 at 600 GB, Inkling-Small BF16 at 600 GB, and Inkling-Small NVFP4 at 180 GB.
The cheapest row still wants a B300.

Straight from the model card: BF16 wants at least 600 GB of aggregated VRAM, so 4 B300s or 8 H200s. The NVFP4 checkpoint wants at least 180 GB, which is W4A4 on one B300, or W4A16 on a pair of H200s.

One GPU. That’s a real milestone for a model this capable, and it’s the first time a Thinking Machines checkpoint fits on a single card at all.

It’s also a card that costs more than a house deposit. Nobody is running this next to their editor. If you want weights on hardware you actually own, the honest answer is still a much smaller model, and our notes on running Qwen 3.7 offline haven’t changed. What Inkling-Small changes is the rental bill: a single-node deployment instead of a four or eight-GPU cluster, and 12B active parameters burning per token instead of 41B. For anyone serving this at volume, that second number is the one that shows up on the invoice.

When we’d reach for it

If you’re already serving open weights for agentic coding and you have retrieval in the loop, Inkling-Small looks like a straight upgrade on cost. Same or better scores on the tasks you’re running, a quarter of the parameters, single-node inference.

If your product answers questions from the model’s own knowledge, keep the flagship, or keep whatever closed model you’re paying for. The SimpleQA gap isn’t a rounding error and no amount of prompting closes it.

And if you’re comparing this against the rest of the open field rather than against its own parent, the picture is less flattering than the launch post suggests. GLM-5.2 still leads on the long-horizon coding benchmarks, and DeepSeek-V4-Flash-0731 posts a much higher Terminal-Bench figure on its own harness. Every one of these numbers is self-reported by the lab that shipped the model, on its own harness, which is the caveat we end up writing every single time.

One thing we can’t check yet: nobody outside the lab has published independent evals of Inkling-Small. Artificial Analysis put it within a point of the flagship on its Intelligence Index, and that’s the only third-party read we found. Give it a fortnight.

Sources

Thinking Machines Lab, Introducing Inkling-Small, 30 July 2026, and the Inkling-Small model card for architecture, hardware and benchmark tables. The flagship figures come from the Inkling model card. Cross-checked against TestingCatalog and XenoSpectrum. Announcement artwork is Thinking Machines Lab’s own, reproduced as a credited citation.

Frequently asked questions

What is Inkling-Small?

Inkling-Small is the second open-weights model from Thinking Machines Lab, released 30 July 2026 under Apache 2.0. It is a 42-layer sparse mixture-of-experts transformer with 276B total parameters and 12B active per token, routing each token to 6 of 256 experts plus 2 shared experts. It takes text, images and audio, returns text only, and supports a context window of up to 1M tokens.

Is Inkling-Small better than the 975B Inkling?

On agentic coding, yes. Thinking Machines' own model cards put Inkling-Small at 80.2% on SWE-bench Verified against 77.6% for Inkling, 64.7% against 63.8% on Terminal-Bench 2.1, and 82.2% against 79.8% on IFBench. It is clearly worse at recall and at long tool-use conversations: SimpleQA Verified falls from 43.9% to 20.6%, and Tau 3 Banking from 23.7% to 15.5%.

What hardware do I need to run Inkling-Small?

The BF16 checkpoint needs at least 600 GB of aggregated VRAM, which the model card lists as 4 Nvidia B300s or 8 H200s. The NVFP4 quantized checkpoint drops that to at least 180 GB, running W4A4 on a single B300 or W4A16 on two H200s. No consumer or workstation GPU gets close, so this is a rented node rather than a desktop model.

Does Inkling-Small still have the thinking effort dial?

Yes. The same controllable thinking effort carries over, exposed as named levels from minimal up to xhigh, so you trade reasoning tokens against latency and cost without swapping models. With 12B active parameters instead of 41B, each of those reasoning tokens is also cheaper to generate.

Where can I get the weights?

Full weights are on Hugging Face under Apache 2.0, with a separate Model Acceptable Use Policy layered on top. Fine-tuning runs on Thinking Machines' Tinker platform, which also hosts a playground chat that accepts text, images and audio.

Tags: aillmnewsopen-sourceself-hostingthinking-machines
Share198Tweet124
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.