PrismML released Bonsai 2 27B on Thursday, a compressed version of Alibaba's open-source Qwen3.8 27B that shrinks the model to 5.9 GB — a 9x to 10x reduction in memory that makes it small enough to fit on a PC and, potentially, a high-end smartphone. The Caltech-founded startup is betting that capable reasoning models don't have to be large, and Bonsai 2 retains 98% of Qwen's aggregate benchmark scores, up from 95% on the first Bonsai model released in March. PrismML has raised a $22.25 million seed round from Khosla Ventures, Cerberus Capital, and Caltech.
The original Bonsai has already been downloaded more than 11 million times on Hugging Face, with PrismML's smaller variants adding another 2.6 million downloads. That traction is what makes the $22.25 million seed look small relative to the company's footprint — Multiverse Computing, a Spanish competitor working on similar compression tech, has raised considerably more. PrismML is led by Babak Hassibi, a Caltech professor who specializes in compression, and counts Ion Stoica, the Databricks co-founder and director of Berkeley's Sky Computing Lab, as an adviser.
The compression technique hinges on how PrismML stores model weights — the numerical parameters a model learns during training. A standard weight requires 16 bits of storage. PrismML uses what it calls ternary weights, collapsing each weight into just three possible values: +1, −1, or 0. With dramatically less information to store per weight, the entire model shrinks by roughly an order of magnitude while preserving nearly all of its behavior.
Key facts
- 01PrismML released Bonsai 2 27B, a compressed version of Alibaba's Qwen3.8 27B that fits in 5.9 GB — a 9x to 10x memory reduction.
- 02Bonsai 2 retains 98% of Qwen's aggregate benchmark scores, up from 95% on the original Bonsai released in March.
- 03The original Bonsai model has been downloaded 11 million times on Hugging Face; smaller PrismML models add another 2.6 million downloads.
- 04PrismML has raised a $22.25 million seed round from Khosla Ventures, Cerberus Capital, and Caltech.
- 05The company plans to release compressed models in the several-hundred-billion-parameter range within the next couple of months.
The two-percentage-point gap between Bonsai 2 and the uncompressed Qwen3.8 27B is largely academic. Benchmarks are imperfect proxies for real-world tasks, and uncompressed frontier models themselves post middling accuracy on many of them. The surrounding software harness a model runs inside typically matters as much as small benchmark deltas. Hassibi has said the goal now is to push compression up the parameter curve.
“The next models that we will release, hopefully in the next couple of months, will be in the several-hundred-billion-parameter range, and I expect it will be easier to retain the intelligence there.”— Babak Hassibi, PrismML CEO and Caltech professor
Counterintuitively, Hassibi argues larger models are easier to compress without losing capability, because they carry more redundant capacity. If that holds, a compressed several-hundred-billion-parameter model could deliver frontier-class reasoning on consumer hardware — the scenario Apple has been chasing with on-device AI. Hassibi declined to comment on rumors that PrismML is in talks with Apple.
The commercial pitch is straightforward: on-device inference cuts cost, latency, and privacy exposure in one move. Every query answered locally is a query the user doesn't have to route to a cloud API, and every model that fits in 5.9 GB is a model a laptop or phone can run without a GPU cluster behind it. That reframes the economics of AI distribution — from metered API calls to software that ships with the device.
Stoica frames the shift in stark terms, arguing that compressed local models change who pays for inference and who sees the data. That framing lands harder as cloud inference bills for enterprises balloon and as regulators in Europe and the US scrutinize how much user data flows to third-party model providers.
The caveats are real. Compression will likely always leave some performance on the table, Hassibi acknowledges, and the 100% parity mark remains unproven at scale. Ternary weights are also not a free lunch — deploying them efficiently requires custom inference kernels and hardware support that isn't yet standard across consumer chips. Multiverse Computing and a growing cohort of compression-focused labs are pursuing overlapping approaches, and it's not clear whose technique wins as models scale.
The strategic implication for the AI market is significant. If a Caltech spin-out with $22.25 million can compress a competitive open-source model 10x with negligible quality loss, the moat around cloud-hosted frontier inference narrows. Every serious model provider — OpenAI, Anthropic, Google — has built revenue models around metered API access to models too large to run locally. PrismML's bet is that most of what those APIs deliver can, within a couple of years, run on hardware users already own. The winners in that scenario are device makers and open-model developers. The losers are the ones charging by the token.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




