COMPANY
PrismML
prismmledge-ai1-bit-llmtopic-note
PrismML
AI startup founded by Caltech researchers, focused on extreme model compression for edge deployment. Emerged from stealth in April 2026 with a $16.25 million seed round and the Bonsai 1-bit LLM family — models where every weight is represented by its sign ({-1, +1}) plus a shared scale factor, trained natively at 1-bit precision end-to-end. The flagship Bonsai 8B fits in 1.15 GB and runs 8x faster than full-precision equivalents on edge hardware. Models released under Apache 2.0.
Timeline
- 2026-04-06-AI-Digest — PrismML emerged from stealth with $16.25M seed round and the open-source Bonsai 1-bit LLM family (8B, 4B, 1.7B) under Apache 2.0, enabling competitive edge inference at a fraction of typical memory requirements.
- 2026-07-15-AI-Digest — PrismML releases Bonsai 27B — 1-bit and ternary quantisations of Qwen3.6-27B that run on-device (HN: 501 pts / 186 cmts). Ternary retains ~95% of FP16 quality across 15 benchmarks; 1-bit retains ~90%. Distinct from the natively-1-bit Bonsai family (April 2026): today’s release applies the compression pipeline to an external open-weight base model rather than training a 1-bit model from scratch. The “27B on a phone” framing needs the quantisation caveat — this is a compression feat, not a new pretrain — but on-device 27B-class inference materially expands what offline LLM assistants can promise.
Context
- Related to MOC - Open Source Models — Bonsai models released under Apache 2.0
- Related to MOC - AI Infrastructure — Extreme compression enabling edge and on-device inference
- Related to Gemma 4 — Competing in the on-device/edge AI space
Timeline
- 2026-07-16-AI-Digest — PrismML ships Bonsai 27B as a fully open reasoning model — a ternary / 1-bit-quantized derivative of Qwen3.6-27B — running on-device on an iPhone. Narrow read: the substantive note is that Bonsai 27B is derivative — a compression story of a Chinese open base rather than an independent open reasoning model, so treat it as more evidence for the Chinese-open-frontier leg of the three-way US/Chinese open/US closed split, not for the US-open below-frontier leg. Structural read: paired with DeepMind‘s verification-bottleneck framing the same digest, PrismML is compounding on the on-device compression axis while frontier labs converge on “the bottleneck is downstream of generation” — two independent AI-for-science-adjacent framings landing the same news cycle.