Skip to main content
Live
Main content

PrismML compresses a 27B model to 5.9 GB with 98% benchmark parity

Bonsai 2 27B shrinks Alibaba's Qwen3.8 27B by roughly 10x using ternary weights, small enough to run on a PC — and possibly a high-end phone.

Jaeden Schafer
Editor in Chief · · 5 min read
PrismML compresses a 27B model to 5.9 GB with 98% benchmark parity

PrismML released Bonsai 2 27B on Thursday, a compressed version of Alibaba's open-source Qwen3.8 27B that shrinks the model to 5.9 GB — a 9x to 10x reduction in memory that makes it small enough to fit on a PC and, potentially, a high-end smartphone. The Caltech-founded startup is betting that capable reasoning models don't have to be large, and Bonsai 2 retains 98% of Qwen's aggregate benchmark scores, up from 95% on the first Bonsai model released in March. PrismML has raised a $22.25 million seed round from Khosla Ventures, Cerberus Capital, and Caltech.

The original Bonsai has already been downloaded more than 11 million times on Hugging Face, with PrismML's smaller variants adding another 2.6 million downloads. That traction is what makes the $22.25 million seed look small relative to the company's footprint — Multiverse Computing, a Spanish competitor working on similar compression tech, has raised considerably more. PrismML is led by Babak Hassibi, a Caltech professor who specializes in compression, and counts Ion Stoica, the Databricks co-founder and director of Berkeley's Sky Computing Lab, as an adviser.

The compression technique hinges on how PrismML stores model weights — the numerical parameters a model learns during training. A standard weight requires 16 bits of storage. PrismML uses what it calls ternary weights, collapsing each weight into just three possible values: +1, −1, or 0. With dramatically less information to store per weight, the entire model shrinks by roughly an order of magnitude while preserving nearly all of its behavior.

Key facts

  • 01PrismML released Bonsai 2 27B, a compressed version of Alibaba's Qwen3.8 27B that fits in 5.9 GB — a 9x to 10x memory reduction.
  • 02Bonsai 2 retains 98% of Qwen's aggregate benchmark scores, up from 95% on the original Bonsai released in March.
  • 03The original Bonsai model has been downloaded 11 million times on Hugging Face; smaller PrismML models add another 2.6 million downloads.
  • 04PrismML has raised a $22.25 million seed round from Khosla Ventures, Cerberus Capital, and Caltech.
  • 05The company plans to release compressed models in the several-hundred-billion-parameter range within the next couple of months.

The two-percentage-point gap between Bonsai 2 and the uncompressed Qwen3.8 27B is largely academic. Benchmarks are imperfect proxies for real-world tasks, and uncompressed frontier models themselves post middling accuracy on many of them. The surrounding software harness a model runs inside typically matters as much as small benchmark deltas. Hassibi has said the goal now is to push compression up the parameter curve.

The next models that we will release, hopefully in the next couple of months, will be in the several-hundred-billion-parameter range, and I expect it will be easier to retain the intelligence there.
Babak Hassibi, PrismML CEO and Caltech professor

Counterintuitively, Hassibi argues larger models are easier to compress without losing capability, because they carry more redundant capacity. If that holds, a compressed several-hundred-billion-parameter model could deliver frontier-class reasoning on consumer hardware — the scenario Apple has been chasing with on-device AI. Hassibi declined to comment on rumors that PrismML is in talks with Apple.

The commercial pitch is straightforward: on-device inference cuts cost, latency, and privacy exposure in one move. Every query answered locally is a query the user doesn't have to route to a cloud API, and every model that fits in 5.9 GB is a model a laptop or phone can run without a GPU cluster behind it. That reframes the economics of AI distribution — from metered API calls to software that ships with the device.

Stoica frames the shift in stark terms, arguing that compressed local models change who pays for inference and who sees the data. That framing lands harder as cloud inference bills for enterprises balloon and as regulators in Europe and the US scrutinize how much user data flows to third-party model providers.

The caveats are real. Compression will likely always leave some performance on the table, Hassibi acknowledges, and the 100% parity mark remains unproven at scale. Ternary weights are also not a free lunch — deploying them efficiently requires custom inference kernels and hardware support that isn't yet standard across consumer chips. Multiverse Computing and a growing cohort of compression-focused labs are pursuing overlapping approaches, and it's not clear whose technique wins as models scale.

Related · from this week
Alibaba releases Qwen3.8-Max, a 2.4-trillion-parameter model rivaling Claude
Jaeden Schafer · 5 min read →

The strategic implication for the AI market is significant. If a Caltech spin-out with $22.25 million can compress a competitive open-source model 10x with negligible quality loss, the moat around cloud-hosted frontier inference narrows. Every serious model provider — OpenAI, Anthropic, Google — has built revenue models around metered API access to models too large to run locally. PrismML's bet is that most of what those APIs deliver can, within a couple of years, run on hardware users already own. The winners in that scenario are device makers and open-model developers. The losers are the ones charging by the token.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Alibaba releases Qwen3.8-Max, a 2.4-trillion-parameter model rivaling Claude
Models

Alibaba releases Qwen3.8-Max, a 2.4-trillion-parameter model rivaling Claude

Alibaba's largest model to date ranks second only to Anthropic's Fable 5 on Arena.AI, with open weights due next week.

Jaeden Schafer5 min read
LLMs believe false claims even when training data labels them as lies
Models

LLMs believe false claims even when training data labels them as lies

A new preprint finds Qwen, Kimi, and GPT-4.1 absorb fabricated facts at an 88.6% belief rate even after explicit negation warnings.

Jaeden Schafer5 min read
Oxford Internet Institute
Models

Oxford study: warmer AI models are 60% more likely to be wrong

Fine-tuning five models including GPT-4o for empathy raised error rates 7.43 points and worsened sharply when users said they felt sad.

Jaeden Schafer5 min read