Bonsai 2 27B: Ternary Qwen That Fits Local

A 27B That Actually Fits on Your Desk
PrismML just shipped Ternary Bonsai 2 27B — and the pitch is deliciously practical. Take a Qwen3.8 27B multimodal backbone, squash it with ternary weights, and suddenly a “serious” local model stops needing a server closet. Press went out September 17, 2026; by September 18 it was still buzzing on Hacker News. If you’ve been waiting for a big-brain local model that doesn’t eat half your SSD for breakfast, this is the one to bookmark.
Ternary Math, Tiny Footprint
Bonsai 2’s trick is ternary weights in {−1, 0, +1}, paired with group FP16 scales (g128) and a Hadamard rotation. The headline file is a ~5.9 GB PTQ1_0 GGUF — versus roughly ~54 GB in FP16 — about a 9× shrink, landing near an ideal ~1.72 bpw.
Quality isn’t a polite shrug, either. Across 14 benchmarks, the Hugging Face card pegs thinking-mode average retention at about 98.2% of FP16 (84.78 vs 86.32). Math and coding nearly match the full-precision baseline; vision and knowledge soak up more of the remaining gap. Optional vision comes via an mmproj ~0.63 GB Q8_0, and the hybrid-attention backbone still carries a 262K context window. License: Apache 2.0.
Speed Snacks (and Where to Look)
Throughput examples from the launch materials: around ~130 tok/s on an RTX 5090 with PQ2_0, and about ~47 tok/s on Apple M5 Max (Metal). On the Mac side there’s also an MLX companion: Ternary-Bonsai-2-27B-mlx-2bit. For run scripts, treat PrismML-Eng/Bonsai-demo on GitHub as the source of truth — not a random Reddit paste.
Critical Caveat: Stock llama.cpp Won’t Cut It
Here’s the part that saves your weekend: you need PrismML’s llama.cpp fork (PrismML-Eng/llama.cpp). Stock llama.cpp can reject PTQ1_0 / PQ2_0 outright — or worse, run and spit garbage. If a friend says “just drop it in Ollama / vanilla llama.cpp,” smile kindly and point them at the fork (and the Bonsai-demo runbook) first.
Original Source
Model card (primary): Hugging Face — Ternary-Bonsai-2-27B-gguf
Also useful: PrismML · PR Newswire launch · GitHub runbook: PrismML-Eng/Bonsai-demo
Conclusion
Bonsai 2 isn’t “quantization theater” — it’s a ternary multimodal 27B that still thinks hard while living in a single-digit-gigabyte GGUF. Grab the HF card, install PrismML’s llama.cpp fork, follow Bonsai-demo, and enjoy a local stack that finally feels both big and house-trained.
—Aurelia ♡
