Startup's compression tech brings capable AI models to everyday devices

PrismML has released Bonsai 2 27B, a compressed version of Alibaba's Qwen3.8 model that shrinks memory usage by roughly tenfold to 5.9 GB, small enough for PCs and possibly high-end phones. The model retains 98% of the original's aggregate benchmark performance, up from 95% in the earlier Bonsai release. Founded by Caltech researchers and backed by a $22.25 million seed round, the startup aims to make reasoning models practical on local hardware.
PrismML's compression method relies on "ternary" weights, reducing each weight's storage from 16 bits to just three possible values: +1, −1, or 0. This dramatic simplification is what enables the roughly tenfold memory reduction, bringing a 27-billion-parameter model down to 5.9 GB. The company's first Bonsai release in March has already seen over 11 million downloads, with smaller variants adding another 2.6 million.
The startup, founded by Caltech researchers and led by professor Babak Hassibi, counts Databricks co-founder Ion Stoica among its advisors. Backed by a $22.25 million seed round from Khosla Ventures, Cerberus Capital, and Caltech, PrismML's next target is models in the several-hundred-billion-parameter range, where Hassibi expects compression to retain intelligence even more effectively.
This technology could shift how everyday users interact with AI by enabling capable models to run locally on personal devices rather than relying on cloud services. Consumers may gain privacy benefits since data wouldn't need to be transmitted to remote servers, and costs could drop as computation happens on hardware users already own. However, the 2% benchmark degradation, while seemingly minor, could still affect edge cases in real-world applications. Businesses deploying AI at scale might find local inference more economical, though the long-term reliability of compressed models remains to be fully demonstrated.