MODEL
Bonsai 27B
modeltopic-noteedge-aiquantizationon-device
Overview
Bonsai 27B is PrismML‘s 1-bit and ternary quantised release of Qwen3.6-27B — a compression feat on an existing model that runs on-device on phones. Ternary quantisation retains ~95% of FP16 quality across 15 benchmarks; 1-bit retains ~90%. Distinct from PrismML’s earlier natively-1-bit Bonsai family (April 2026); this July release applies the compression pipeline to an external open-weight base model rather than training a 1-bit model from scratch. The “27B on a phone” framing needs the quantisation caveat, but on-device inference at 27B-class weights — even lossy — meaningfully expands what offline LLM assistants can promise.
Timeline
- 2026-07-15-AI-Digest — PrismML releases Bonsai 27B as 1-bit and ternary quantisations of Qwen3.6-27B that run on-device (Hacker News: 501 pts / 186 cmts). Ternary retains ~95% of FP16 quality across 15 benchmarks; 1-bit ~90%. Positioned as “a 27B-class model that runs on a phone.”
Key Developments
- On-Device 27B-Class Inference via Quantisation of Qwen3.6-27B (July 15, 2026): 1-bit and ternary quantisations of Qwen3.6-27B — ternary preserves ~95% of FP16 quality across 15 benchmarks, 1-bit ~90%. A compression feat applied to an existing open-weight model, not a native-1-bit pretrain; but on-device inference at 27B-class weights, even lossy, materially expands what offline LLM assistants can offer.
- 2026-07-16-AI-Digest — PrismML ships Bonsai 27B as a fully open reasoning model — ternary / 1-bit-quantised derivative of Qwen3.6-27B running on-device on an iPhone. Reprised as day-two coverage of the July 15 announcement now framed by The Decoder as the on-phone reasoning model landmark. Narrow read: the Bonsai note that matters is derivative — compression story of a Chinese open base, not an independent open reasoning model — so treat it as more evidence for the Chinese-open-frontier leg of the three-way US/Chinese open/US closed split, not the US-open below-frontier leg. Structural read: paired with DeepMind‘s verification-bottleneck framing the same digest, Bonsai’s on-device compression axis and DeepMind’s downstream-of-generation axis are two adjacent AI-for-science-relevant framings landing the same news cycle.
Related
See also: PrismML, Qwen3.6-27B, MOC - Open Source Models.