Artificial Intelligence Startups

The Lab Shrinking Reasoning Models Small Enough to Live on Your Phone

A small research lab out of Caltech has squeezed a 27-billion-parameter reasoning model into 5.9 GB while holding on to 98% of its benchmark performance — a result that could move serious AI off the cloud and onto hardware people already own.

5.9 GB
Compressed size of a 27B-parameter reasoning model
98%
Of the original model's aggregate benchmark scores retained
$22.25M
Seed funding raised so far by the lab

For three years the AI industry has run on a simple assumption: the smarter you want a model to be, the bigger it has to get — and the bigger it gets, the further away it has to live, humming inside someone else's data centre. A young lab called PrismML is quietly arguing that assumption is wrong, and it now has a model small enough to make the argument stick.

9–10x
Memory reduction versus the uncompressed model
13.6M+
Combined downloads across the lab's model family
3
Possible values per weight: +1, −1 or 0

A Small Lab With an Outsized Claim

PrismML has not raised the kind of money that buys headlines. Its seed round sits at $22.25 million — a rounding error next to the war chests of the frontier labs. What it has instead is a specific technical thesis and the credentials to pursue it. The company was started by a group of Caltech researchers and is run by Babak Hassibi, a Caltech professor whose academic career has centred on compression. Ion Stoica, a co-founder of Databricks and director of Berkeley's Sky Computing Lab, advises the company. Khosla Ventures, Cerberus Capital and Caltech are on the cap table.

Its latest release, Bonsai 2 27B, takes Qwen3.8 27B — a widely used open-source model from Alibaba — and compresses it down to 5.9 GB. That is a nine- to ten-fold cut in memory footprint, and small enough to sit comfortably on a laptop and, plausibly, on a high-end phone. Speculation that the lab is in conversation with Apple has been circulating; the CEO has not commented on it.

The Number That Actually Matters

Shrinking a model is not hard. Shrinking one without lobotomising it is. Plenty of compression techniques trade away accuracy until the resulting model is fast, tiny and not worth using. PrismML's pitch rests on how little it appears to give up: Bonsai 2 lands at 98% of the original model's aggregate benchmark scores, up from 95% for the first Bonsai release. Two releases, three points of recovered performance — a trajectory that matters more than either number on its own.

Benchmark Retention Across Compression Releases
Aggregate benchmark performance of the compressed model as a share of the original
0% 25% 50% 75% 100% 95% 98% 100% First Bonsai release Bonsai 2 27B Full parity goal
Measured benchmark retention
Target parity, not yet reached
Source: Startup360hub analysis.

Hassibi is candid that some loss is probably permanent — compression will always cost something. But the remaining gap is arguably academic. Uncompressed models are themselves imperfect, and benchmark scores are a rough proxy for real-world usefulness at best. A two-point difference on a leaderboard rarely translates into a difference a user would notice. The software harness a model runs inside has repeatedly been shown to influence accuracy at least as much as the weights themselves.

How You Fit a Giant Model in a Phone

The technique comes down to how the model stores what it knows. A model's weights are the values it absorbs during training, and conventionally each one is stored using 16 bits of precision. PrismML's approach, which it calls ternary weights, collapses that range to just three possible states: +1, −1 or 0. Storing three options instead of thousands of gradations means each weight consumes a fraction of the space, and the total footprint falls off a cliff.

The obvious objection is that throwing away that much precision should wreck the model. The lab's results suggest it does not — at least not by much — and that is the whole commercial proposition.

📱
Intelligence that ships with the device
A reasoning model that runs locally needs no subscription and no server call. The compute is hardware the user already paid for.
🔒
Privacy as a side effect
If the prompt never leaves the handset, there is no cloud log, no transit and no third-party retention to worry about.
⚡
A different cost curve
On-device inference sidesteps the per-token economics that make consumer AI products structurally unprofitable at scale.
🏁
A crowded starting line
Compression is not an uncontested field. Multiverse Computing, among others, is chasing the same prize with considerably more capital.

What Comes Next

The lab's next target is far larger: models in the several-hundred-billion-parameter range. Counterintuitively, Hassibi expects that to be the easier problem. Bigger models carry more redundancy, which leaves more slack to squeeze out before intelligence starts to degrade. If that holds, full performance parity becomes more achievable as models grow, not less — the opposite of what the scaling-is-everything narrative would predict.

Adoption suggests developers are already convinced enough to try it. The first Bonsai model has been downloaded more than 11 million times, with a further 2.6 million downloads across the lab's smaller variants. For a company most people have never heard of, that is meaningful distribution.

The strategic question for founders is what a world of capable local models does to the assumptions baked into today's AI products. A great deal of current infrastructure — API billing, usage tiers, inference margin, cloud lock-in — exists because intelligence has to be rented. If a sufficiently good model fits on the device in a buyer's pocket, that whole layer becomes optional for a large class of applications.

The most disruptive thing in AI right now may not be a bigger model at all. It may be a small one that works well enough to make the data centre unnecessary.

— Startup360hub
🔑 Key Takeaways
  1. A 27B reasoning model now fits in 5.9 GB. That is a nine- to ten-fold memory reduction, small enough for a laptop and potentially a premium smartphone.
  2. Performance loss is marginal and shrinking. Benchmark retention has moved from 95% to 98% across two releases, with the remaining gap unlikely to be noticeable in practical use.
  3. Ternary weights are the mechanism. Reducing each weight from 16 bits of precision to three states — +1, −1 or 0 — is what collapses the footprint.
  4. Bigger models may compress better, not worse. The lab expects several-hundred-billion-parameter models to retain more of their intelligence because they carry more redundancy.
  5. The business-model implications are the real story. Free, private, offline inference on hardware users already own undercuts the rental economics most AI products are built on.
Topics Artificial Intelligence Model Compression On-Device AI Open Source LLMs Deep Tech Startups