Apple On-Device AI Talks: PrismML Says It Can Squeeze a 27B Model Onto an iPhone
Apple’s push toward on-device AI just got a credible technical assist: PrismML, a Caltech spinout backed by Khosla Ventures, confirmed on July 14 that it is in talks with Apple about compression technology that shrinks server-class models down to phone size. CEO Babak Hassibi told CNBC that Apple is “really evaluating our technology right now,” describing the discussions as very early but “progressing nicely.”
The Numbers PrismML Is Claiming
The demo that got Apple’s attention involves Alibaba’s open-source Qwen model. In its standard form, the 27-billion-parameter version weighs about 54 gigabytes — usable only on a Mac with 64GB of memory or more. PrismML says its compression brings the same model under 4 gigabytes, which fits comfortably inside the 8GB of RAM in an iPhone 15 Pro.
The broader claims:
- 10–15x reduction in memory usage
- 6–8x increase in processing speed
- 3–6x lower energy consumption
- A loss of only “a few percentage points” of overall performance
That last bullet deserves scrutiny, and to the company’s credit Hassibi did not hide it. The compressed models are measurably weaker at factual reasoning, mathematics, and coding — which happens to be exactly the cluster of tasks where users notice degradation fastest. A shrunken model that handles summarization and rewriting well but fumbles arithmetic is a different product than a full-size one.
Why On-Device AI Matters So Much to Apple
Apple has built its AI positioning around privacy, and privacy is easiest to guarantee when the data never leaves the phone. The constraint has always been physics: frontier models do not fit in a pocket, so anything genuinely capable has been routed to Private Cloud Compute or, for harder queries, to third-party models entirely.
Effective on-device AI would collapse that split. It would mean no internet dependency, no server round-trip latency, and no need to explain to regulators where a query went. For a company that has spent two years apologizing for Siri, a 27B-class model running locally is the difference between a marketing claim and a shipping feature.
The Quiet Implication for the Memory Crisis
There is a second-order story here that is arguably bigger than Siri. The AI infrastructure buildout has consumed global supply lines for memory and chips, and that shortage is what is currently driving up hardware prices across the industry — including, by Apple’s own admission, the iPhone.
If meaningful inference moves onto devices people already own, demand for cloud inference capacity softens at the margin. Nobody should expect a 27B model on an iPhone to unwind a multi-hundred-billion-dollar data center cycle. But the direction of the arrow matters: every workload that runs locally is a workload not competing for HBM. AppleInsider’s report on the confirmed talks makes the same connection to the component crunch.
Licensing or Acquisition?
Worth noting how this became public at all. Apple acquires companies in near-total silence — the news usually breaks when someone’s LinkedIn changes. The fact that PrismML’s CEO is talking to CNBC about Apple’s interest is itself evidence against an acquisition being close, since an acquisition target would almost certainly be under NDA.
A licensing deal is the more plausible read. It also puts PrismML in the awkward but enviable position of having advertised itself to every one of Apple’s competitors at the same time. Hassibi confirmed other companies are evaluating the technology too.
The Bottom Line
Treat the compression numbers as vendor claims until independent benchmarks land — the performance loss on reasoning and coding is the figure to watch, not the 54GB-to-4GB headline. But the strategic logic is sound, and it explains why Apple is looking. On-device AI has been Apple’s stated destination since Apple Intelligence launched. Someone finally showed up with a plausible route.
Related on DAILYSIM: iPhone Price Increase Confirmed: Apple Blames the AI Memory Shortage and New York Data Center Ban: Hochul Freezes Hyperscale AI Projects for a Year.