PrismML hopes its tiny LLM will change how we all use AI — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- PrismML released Bonsai 2 27B on Thursday, compressing Alibaba's Qwen3.8 27B open-source model from a much larger size down to 5.9 GB — a 9x to 10x memory reduction that fits on a PC and possibly a high-end smartphone
- Bonsai 2 retains 98% of Qwen's aggregate benchmark scores, up from 95% in the first Bonsai released in March; the original Bonsai has been downloaded over 11 million times, with smaller models downloaded another 2.6 million times
- PrismML's compression works by reducing model weights — the information learned during training — from the standard 16 bits down to ternary values of +1, −1, or 0
- CEO Babak Hassibi, a Caltech professor and compression expert, said the startup plans to apply the technique to several-hundred-billion-parameter models within the next couple of months
- PrismML has raised just a $22.25 million seed round from Khosla Ventures, Cerberus Capital, and Caltech, with Ion Stoica — Databricks co-founder and Berkeley Sky Computing Lab director — serving as an adviser
- Adviser Ion Stoica framed the impact as 'intelligence at your fingertips… free because it's going to run on the device you already bought… private, because you're not going to send it to the cloud'
- Competitor Multiverse Computing, founded by a Spanish physics professor, is pursuing the same LLM-compression space but has raised significantly more capital; PrismML is rumored to be in talks with Apple, though Hassibi declined to comment
Why it matters: PrismML's 9-10x compression with only 2% benchmark degradation challenges the assumption that capable reasoning models require massive cloud infrastructure — and Hassibi argues larger models are actually easier to compress without losing intelligence, with several-hundred-billion-parameter versions targeted within months. If the ternary-weight approach holds at scale, the AI cost structure could shift from cloud inference to on-device execution, a shift Stoica says delivers both privacy and zero marginal cost to users.
Ask SkimNews

