Squeezing a reasoning model small enough to run on a phone usually costs accuracy. PrismML says the bill is now 2 percent.
The Caltech spinout released Bonsai 2 27B on Thursday. It compresses Qwen3.8 27B, a popular open source model from Alibaba, to 5.9GB, roughly a 9x to 10x memory reduction over the original and small enough for a PC and possibly a high-end handset.
The technique is ternary quantization. Instead of the 16 bits a weight normally needs, each weight becomes one of three values: +1, -1 or 0. Storing fewer possible values per weight is what shrinks the file.
PrismML reports that the compressed model holds 98 percent of Qwen’s aggregate benchmark scores. The first Bonsai, out in March, held 95 percent, so the gap is narrowing release by release. That earlier model has been downloaded more than 11 million times, and the company’s smaller models another 2.6 million times.
The lab came out of Caltech and is led by CEO Babak Hassibi, a professor there who works on compression. Ion Stoica, a Databricks co-founder who directs Berkeley’s Sky Computing Lab, advises. Khosla Ventures, Cerberus Capital and Caltech backed the $22.25M seed.
Hassibi would not discuss rumors of talks with Apple, and he cautioned that compression will always cost something, so perfect parity is unlikely to arrive. A 2 percent gap may not matter much in practice anyway, since accuracy depends heavily on the harness a model runs inside.