<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/bonsai-27b-a-27b-multimodal-model-that-runs-on-a-phone-at-3-9gb-tzzuutsat" -->

---
title: Bonsai 27B: a 27B multimodal model that runs on a phone...
description: PrismML has announced Bonsai 27B, a 27-billion-parameter multimodal model optimized to run directly on mobile devices using only 3.9GB of memory. The company...
canonical: https://daily.dev/posts/bonsai-27b-a-27b-multimodal-model-that-runs-on-a-phone-at-3-9gb-tzzuutsat
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Bonsai 27B: a 27B multimodal model that runs on a phone at 3.9GB | daily.dev
og:description: PrismML has announced Bonsai 27B, a 27-billion-parameter multimodal model optimized to run directly on mobile devices using only 3.9GB of memory. The company...
og:url: https://daily.dev/posts/bonsai-27b-a-27b-multimodal-model-that-runs-on-a-phone-at-3-9gb-tzzuutsat
og:image: https://api.daily.dev/og/posts/tzzuUtSAT.png
og:image:alt: Bonsai 27B: a 27B multimodal model that runs on a phone at 3.9GB
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Bonsai 27B: a 27B multimodal model that runs on a phone at 3.9GB

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 2 upvotes · 0 comments

## Summary

PrismML has announced Bonsai 27B, a 27-billion-parameter multimodal model optimized to run directly on mobile devices using only 3.9GB of memory. The company claims it is the first 27B-class model capable of on-device inference, though no further technical details about the compression techniques used were provided in the announcement.

## Content

PrismML released Bonsai 27B, a multimodal model that fits on an iPhone, Android device, or Mac. That's a 27-billion parameter model in 3.9GB, made possible through 1-bit and ternary quantization that the team claims retains about 95% of the quality of a full fp16 model.

The speed numbers are real: on an NVIDIA RTX 5090, it hits up to 163 tokens/sec in 1-bit and 134 tok/s in ternary. On an M5 Max, you're looking at 87 tok/s and 58 tok/s respectively. For on-device inference, those are genuinely fast.

The model is based on Qwen and sits roughly at the capability level of Claude Sonnet from about six months ago. So no, it's not going to replace your cloud API for serious work. But that's not really the point.

Two years ago, running a model this size required a server room. Now it runs in a browser via WebGPU (the webml-community has a Hugging Face Space for exactly that), on your phone, or through Claude Code via Hugging Face's integration with Together Compute.

The trajectory here matters more than the current benchmark position. Local models have gone from multi-million dollar data center territory to fitting in your pocket in roughly 24 months. If that curve continues - and there's no obvious reason it won't - the models running on a Mac Mini in 12 months will be meaningfully better than what's available through cloud APIs today.

Ternary-Bonsai-27B is available on Hugging Face under the prism-ml organization.

## Similar posts on daily.dev

- [I ran the tiny Bonsai model on my tiny GPU. Here’s how it performed](https://daily.dev/posts/i-ran-the-tiny-bonsai-model-on-my-tiny-gpu-here-s-how-it-performed-dsp1h8gia) · InfoWorld · 0 upvotes · 0 comments
- [PrismML — Introducing Ternary Bonsai: Top Intelligence at 1.58 Bits](https://daily.dev/posts/prismml-introducing-ternary-bonsai-top-intelligence-at-1-58-bits-klkxnwryn) · Hacker News · 1 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#multimodal](https://daily.dev/tags/multimodal), [#local-ai](https://daily.dev/tags/local-ai)

[View this post on daily.dev](https://daily.dev/posts/bonsai-27b-a-27b-multimodal-model-that-runs-on-a-phone-at-3-9gb-tzzuutsat)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Bonsai 27B: a 27B multimodal model that runs on a phone at 3.9GB","url":"https://daily.dev/posts/bonsai-27b-a-27b-multimodal-model-that-runs-on-a-phone-at-3-9gb-tzzuutsat","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/bonsai-27b-a-27b-multimodal-model-that-runs-on-a-phone-at-3-9gb-tzzuutsat"},"datePublished":"2026-07-14T18:24:38.395Z","dateModified":"2026-07-20T09:38:29.586Z","description":"PrismML has announced Bonsai 27B, a 27-billion-parameter multimodal model optimized to run directly on mobile devices using only 3.9GB of memory. The company...","image":"https://pbs.twimg.com/media/HNNUKpZXgAAuAKd.jpg","thumbnailUrl":"https://pbs.twimg.com/media/HNNUKpZXgAAuAKd.jpg","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/bonsai-27b-a-27b-multimodal-model-that-runs-on-a-phone-at-3-9gb-tzzuutsat","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,multimodal,local-ai","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Bonsai 27B: a 27B multimodal model that runs on a phone at 3.9GB"}]}
```

