• Latest
  • Trending
  • All
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

7 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Sunday, September 20, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

by stephane
20 September 2026
in Dev
0
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.
491
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

We've got a 24 GB card in the office box and a 27B model we'd like to run on it with room left for a real context. Until this week the honest options were a 4-bit GGUF at 17.6 GB or a "2-bit" one that scored like an 8B. On 17 September 2026 PrismML released Ternary Bonsai 2 27B, a version of Qwen3.8 27B where every weight is -1, 0 or +1, and the file is 5.95 GB. Apache 2.0, 262K context, vision tower optional. The headline is 98.2% of the full-precision score, and that number is real. So is the one they keep outside the average.

The short answer

Ternary Bonsai 2 27B is Qwen3.8 27B squeezed to 1.72 bits per weight, shipped as two GGUF packs (5.95 GB and 7.21 GB) plus an MLX pack for Apple silicon. On PrismML's 14-benchmark thinking-mode suite it averages 84.78 against 86.32 for FP16, within 0.4 points of the 17.6 GB 4-bit build and 12 points above a conventional 2-bit one. On the two long-horizon agent benchmarks the whitepaper keeps out of that average, it keeps about three quarters: 60.8 on SWE-bench Verified where the parent scores 80.6. You can't load it in stock llama.cpp; it needs PrismML's fork. Every number here is PrismML's own, and nobody's reproduced them yet.

5.95 GBfor 27.36B parameters, against 53.8 GB in FP16
84.78average on 14 benchmarks, FP16 scores 86.32
60.8on SWE-bench Verified, the parent gets 80.6
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, keeps about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.
Nine times smaller than FP16 and one benchmark suite where it holds. The other suite is the story.

What 1.72 bits per weight buys you

Ternary means each weight stores one of three values, and log2(3) is about 1.585 bits. Add one FP16 scale shared by every group of 128 weights and you land at 1.71. A sliver of tensors (26.2 million parameters, under 0.1%) stays in higher precision for the recurrent state path and the norms, which brings the whole model to 1.72. Real kernels need a packed layout, so there are two GGUF files. PTQ1_0 packs the trits densely at 1.75 bits per weight and 5.95 GB. PQ2_0 gives each trit a 2-bit slot, costs 7.21 GB, and is cheaper to unpack. The optional vision tower is a separate 0.63 GB mmproj in Q8_0 that only loads when you send an image.

That's the weights. Not the memory. The KV cache for a 262K context sits on top, and the base is a hybrid with roughly 75% linear-attention layers, which is what makes long context survivable at all on a laptop. PrismML's own line is "a 16 GB laptop or a single 24 GB GPU", and honestly that matches what we'd expect: a 7.21 GB file plus a 32K context fits a 4090 with plenty spare, and the Apple numbers were measured on M5 Pro and M5 Max machines running the 7.2 GB pack. The base is the same Qwen3.8-27B that shipped under Apache 2.0 in August, so the licence carries through cleanly.

Speed is bandwidth-bound at batch 1, and the model card's llama-bench table is refreshingly specific. RTX 5090: 129.9 tokens a second on PQ2_0, 120.5 on PTQ1_0. RTX 4090: 81.2 and 91.1. H100: 113.9 and 86.9. An L4 at 72 W: about 30 either way. M5 Pro laptop: 28.1 tokens a second on Metal, drawing 27.5 W on the GPU rail, against 300 to 455 W of board power on the Nvidia cards. The launch post says "up to 143 tokens a second" on the 5090, which is higher than anything in the model card table, and both numbers are theirs. I'd plan around the table. The smaller pack isn't always the faster one either: dense trits cost arithmetic to unpack, so PTQ1_0 wins on the 4090, the Ada 6000 and the L4, and loses on Hopper, Ampere and Blackwell where batch-1 decode isn't starved for memory. On a 5090 or an H100, download the 7.21 GB file.

Two benchmark suites, and only one in the headline

The 98.2% comes from an average over 14 thinking-mode benchmarks on the model card, or 20 in the whitepaper. Same retention figure both ways. The card's version reads: FP16 86.32, the 4-bit UD-Q4_K_XL build at 17.6 GB scores 85.18, a conventional IQ2_XXS at 9.4 GB scores 72.59, and Bonsai 2 at 5.9 GB scores 84.78. The per-benchmark rows show where the conventional 2-bit build dies: 57.5 on AIME26 and 56.4 on LiveCodeBench, while still posting 88.93 on MMLU-Redux, which is exactly why a quick chat test wouldn't catch it. Bonsai 2 posts 95.83 and 90.07 on those two, level with or above FP16. Math loses half a point (96.57 against 97.06 by category). Coding is flat. Instruction following is actually ahead.

Horizontal bar chart of average scores on PrismML's 14 thinking-mode benchmarks for four builds of Qwen3.8 27B, showing FP16 at 86.32 and 54 gigabytes, the 4-bit UD-Q4_K_XL build at 85.18 and 17.6 gigabytes, Ternary Bonsai 2 27B at 84.78 and 5.9 gigabytes, and the conventional 2-bit IQ2_XXS build at 72.59 and 9.4 gigabytes.
Same base model, four files. The 9.4 GB one is the trap, and the 5.9 GB one lands next to the 17.6 GB one.

Where it drops is the knowledge and vision rows. MuSR falls from 79.63 to 70.63, MMLU-Redux from 91.46 to 89.09, MMMU-Pro from 81.73 to 75.49 and OCR Bench v2 from 60.99 to 56.88. That pattern (reasoning intact, recall and perception softer) is consistent with a model that lost bits rather than structure, and it's the same shape PrismML reported on Bonsai 1 two months ago, when the ternary variant kept about 95%.

Then there's the pair the whitepaper reports but leaves out of the average, and MarkTechPost was right to pull them up. Terminal-Bench 2.1: 52.8 against 69.7 for the parent. SWE-bench Verified: 60.8 against 80.6. That's about 75% retained, not 98%. Long-horizon agent tasks compound small errors across dozens of tool calls, so a model that's 2% worse per step doesn't come out 2% worse at the end. PrismML's launch copy talks up "long-horizon agentic performance" and demos the model driving a coding agent on a 5090. Both things can be true. PrismML says it's better than Bonsai 1 at this, and it's still 20 points behind the model it was made from. For a local coding agent, that's the number to weigh, and it's the one we'd want to see someone outside PrismML measure. One tester on GitHub who ran the v1 model through an agent suite wrote that the claimed lead "did not hold at suite level" on wall time, and has a v2 run queued. We'd read that before anything else.

Running it, which means a fork

Stock llama.cpp rejects PTQ1_0 and PQ2_0 as unknown tensor types. Worse, the card warns it'll load a Q2_0 file without complaint and produce garbage, because the ternary path relies on a Hadamard rotation applied at runtime that mainline doesn't have. So you build or download PrismML-Eng/llama.cpp, and on Apple you use their MLX fork instead. There's a prebuilt release archive per platform if you'd rather not compile.

bash
git clone https://github.com/PrismML-Eng/llama.cpp && cd llama.cpp && cmake -B build -DGGML_CUDA=ON && cmake --build build -j
bash
hf download prism-ml/Ternary-Bonsai-2-27B-gguf Ternary-Bonsai-2-27B-PQ2_0.gguf --local-dir . && ./build/bin/llama-cli -m Ternary-Bonsai-2-27B-PQ2_0.gguf -ngl 99 -fa on -c 32768 --temp 1.0 --top-p 0.95 --top-k 20 -p "Explain the difference between a /24 and a /25." -n 256

Drop -DGGML_CUDA=ON on macOS, Metal is the default there. -ngl 99 puts every layer on the GPU and -c goes up to 262144 if you've got the memory for the cache. It's a reasoning model and it thinks by default, so budget output tokens accordingly. The sampling flags are the ones PrismML recommends, and the repo they call the source of truth is PrismML-Eng/Bonsai-demo, which pins a known-good binary and covers the server, tool calling and image input with the mmproj file.

The fork is the part we'd flag to anyone running a fleet. A vendor fork of llama.cpp means you're on their release cadence for kernel fixes, and every upstream feature lands on their schedule or not at all. That's fine for a workstation. For a serving box it's a dependency you didn't have with a standard GGUF of a Qwen model. PrismML says returning the footprint advantage as latency on every GPU is "an active engineering target", which is a candid way of saying the kernels aren't done. I might be wrong about how long that takes, but ternary kernels in mainline llama.cpp would change the calculus completely, and nothing on either repo says it's coming.

Checklist of what Ternary Bonsai 2 27B ships and what it does not as of 20 September 2026, listing Apache 2.0 weights in 5.95 and 7.21 gigabyte GGUF packs plus an MLX pack, a 262K context with an optional 0.63 gigabyte vision tower, 84.78 against 86.32 on the 14-benchmark average, and on the other side that stock llama.cpp cannot load it, that SWE-bench Verified and Terminal-Bench 2.1 keep only about 75 percent, and that no independent reproduction exists yet.
Three rows from the model card, three from the whitepaper and the runtime notes. All of them PrismML's own words.

Sources

PrismML, Introducing Bonsai 2 27B, 17 September 2026 (the release date, the 98.2% and 83.9 against 85.4 figures, the "up to 143 tokens/second" claim, Apache 2.0, the 95% retention of Bonsai 1). Hugging Face, prism-ml/Ternary-Bonsai-2-27B-gguf model card, read 20 September 2026 (the 1.72, 1.75 and 2.13 bits per weight, the 5.95 and 7.21 GB packs, the 0.63 GB vision tower, the 14-benchmark table and per-benchmark rows, the llama-bench throughput and power table, the fork requirement and the run commands). PrismML, Bonsai 2 27B whitepaper and PrismML-Eng/llama.cpp, September 2026 (the 20-benchmark average, the 26.2M high-precision parameters, the packings). MarkTechPost, PrismML releases Ternary Bonsai 2 27B, 18 September 2026 (the Terminal-Bench 2.1 and SWE-bench Verified rows from the whitepaper, the "16 GB laptop or a single 24 GB GPU" line, the Cline demo). GitHub, evanwtf/local-llm issue 479, September 2026 (an independent tester's note on the v1 agent results and a planned v2 run).

Frequently asked questions

How much VRAM does Ternary Bonsai 2 27B need?

The weights are 5.95 GB (PTQ1_0) or 7.21 GB (PQ2_0), plus 0.63 GB if you load the vision tower. The KV cache comes on top and scales with the context you set. PrismML's guidance is a 16 GB laptop or a single 24 GB GPU; their own measurements ran on a 24 GB RTX 4090, a 24 GB L4 and M5 Pro and M5 Max laptops.

Does it run in Ollama, LM Studio or regular llama.cpp?

Not as of 20 September 2026. The PTQ1_0 and PQ2_0 tensor types and the runtime Hadamard rotation only exist in PrismML's llama.cpp fork (and their MLX fork for Apple silicon). Stock llama.cpp rejects the files, and the model card warns that a Q2_0 file loads silently and produces garbage. Anything built on mainline llama.cpp inherits the same limit until the kernels are upstreamed, and nothing says they will be.

Is 98.2% the right number to quote?

It's the average over PrismML's 14 (model card) or 20 (whitepaper) thinking-mode benchmarks, and it's accurate for that suite. It doesn't cover Terminal-Bench 2.1 or SWE-bench Verified, where the whitepaper reports 52.8 against 69.7 and 60.8 against 80.6, about 75% retained. If your use is a long-running coding agent, quote 75%.

What's the licence?

Apache 2.0, inherited from Qwen3.8-27B. That covers the weights on Hugging Face. The llama.cpp fork carries llama.cpp's MIT licence; check the fork's own LICENSE file for anything PrismML added.

Which file should we download, PTQ1_0 or PQ2_0?

PQ2_0 (7.21 GB) unless disk or memory is the constraint. It's faster on the RTX 5090, H100, A100 and the Blackwell workstation cards, and prompt processing favours it everywhere. PTQ1_0 (5.95 GB) is faster on the RTX 4090, the RTX 6000 Ada, the L40S and the L4, where memory bandwidth is the bottleneck.

Tags: Bonsai 2llama.cpplocal LLMnewsopen-weightsPrismMLquantizationQwen3.8
Share196Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.