Skip to content
AIForker

AI Tools, Tutorials, and Insights。

AIForker

AI Tools, Tutorials, and Insights。

  • Home
  • AI Tool Reviews
  • AI Guides
  • AI Agent
    • Codex
    • Hermes
    • Openclaw
    • Claude Code
    • Gemini
  • China AI
    • DeepSeek
    • GLM
    • Qwen
    • Doubao
    • MiniMax
    • Seedance
    • Kimi‌
    • iFLYTEK Spark
  • AI Prompts
  • About Us
  • Home
  • AI Tool Reviews
  • AI Guides
  • AI Agent
    • Codex
    • Hermes
    • Openclaw
    • Claude Code
    • Gemini
  • China AI
    • DeepSeek
    • GLM
    • Qwen
    • Doubao
    • MiniMax
    • Seedance
    • Kimi‌
    • iFLYTEK Spark
  • AI Prompts
  • About Us
  • https://www.facebook.com/
  • https://twitter.com/
  • https://t.me/
  • https://www.instagram.com/
  • https://youtube.com/
Home/China AI/Qwen/27B on a Phone: The Density Numbers Behind PrismML’s Bonsai Breakthrough
QwenAI Guides

27B on a Phone: The Density Numbers Behind PrismML’s Bonsai Breakthrough

By Forker
July 14, 2026 3 Min Read
0

A 27B model that runs on a phone used to be a research talking point. PrismML’s Bonsai 27B makes it real. The 1-bit variant is 5.9GB. That is the headline number, and it is backed by a benchmark comparison that is worth walking through in detail.

What Bonsai 27B Actually Is

Bonsai 27B is based on Qwen 3.6 27B, a known strong performer in the 27B class. PrismML did not redesign the architecture. They changed how the weights are stored. Ternary weights from {-1, 0, +1} with FP16 group-wise scaling give 1.71 effective bits per weight. The 16-bit original needs 54GB. Bonsai 1-bit needs 5.9GB. The ternary variant sits at 18GB for laptops with discrete GPUs.

9x compression. That is the underlying number everything else rests on.

The Density Numbers

PrismML evaluates this on intelligence density: benchmark score divided by model size in GB. Full-precision Qwen 3.6 27B scores about 0.05 per GB. Bonsai 1-bit scores 0.53 per GB. Ten times the density.

For local inference, that is a different cost structure. No API calls. No per-token billing. Data stays on device. If your use case requires privacy, offline operation, or sub-100ms response, the math shifts in favor of local in a way it did not twelve months ago.

The Benchmark Gaps

Ternary Bonsai 27B vs full-precision Qwen 3.6 27B:

– Math: 93.4 vs 95.3, gap 1.9
– Coding: 86.0 vs 88.7, gap 2.7
– Agentic/Tool-calling: 74.0 vs 80.0, gap 6.0
– Instruction following: 71.8 vs 78.4, gap 6.6
– Knowledge/STEM: 77.0 vs 83.1, gap 6.1
– Vision: 65.2 vs 72.6, gap 7.4
– Overall: 80.5 vs 85.0, gap 4.5

Math and coding show the smallest gaps. Vision and knowledge show the widest. The overall gap is 4.5 points. For tasks that do not push the edges of vision or knowledge retrieval, the practical difference in everyday use is small.

What the Model Can Actually Do

Bonsai 27B ships with multi-step reasoning, structured tool calls, vision input, and computer-use agent loops. Demo videos show it running the Hermes agent. These are not features added after quantization. They are part of the core capability set, which makes this more of a production bet than a research demo.

The Tradeoff Is Real

Ternary quantization is lossy. Group-wise FP16 scaling recovers some of what is lost when weights are forced into three discrete values, but not all of it. For knowledge-intensive tasks or vision tasks that need high fidelity, full-precision still wins. For most consumer-facing features, the gap is small enough that users will not notice in practice.

One thing worth flagging: PrismML published a whitepaper on the quantization method. That is not standard in this space. Most quantized releases come with marketing claims and no methodology to audit. Having something to point to and verify is the exception, not the rule.

What This Means for Mobile AI

The 18GB ternary variant is the practical entry point for most developers right now. Laptops with 16GB of RAM can handle it without issues. The 5.9GB 1-bit variant is the milestone, but iOS and Android deployment requires inference stack work that is still maturing.

The 10x density number is the thing I keep coming back to. 54GB to 5.9GB is not incremental compression. It moves the model from a category that requires datacenter hardware to a category that fits in your pocket. Whether the benchmark gaps matter depends on what you are building. For coding and reasoning tasks, which drive most of the actual demand for large models, the gaps are small. For vision-heavy or knowledge-heavy tasks, they are still there and worth accounting for.

PrismML has put out a whitepaper and a set of benchmark numbers. The next step is independent testing. If the density numbers hold up outside PrismML’s own evaluation environment, this is a meaningful step forward for on-device AI.

Related Articles:

  1. Un-0: The AI Image Generator That Uses 1000x Less Energy Than GPUs
  2. The world’s first open-source MoE video-based model, LingBot-Video for embodied intelligence
  3. GPT-5.6 Is Here — Same Price as GPT-5.5, Twice the Brain
  4. The 5 Tools That Actually Make Hermes Worth Using
  5. Interactive Animations With Codex: No Design Skills Required
  6. Your AI Coding Assistant Has No Memory. Here’s Why That Matters

Tags:

local-aiBonsai-27BPrismMLon-device-AIefficient-AIternary-quantizationedge-AIsmall-AI-model
Author

Forker

Follow Me
Other Articles
Previous

Tencent Hunyuan Hy3 Quantized Models: 295B Parameters, Single-GPU Deployment

SafeLine
Next

This Open-Source WAF from China Has 21.7K GitHub Stars

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest Articles

  • Codex + OpenMontage Made Me Throw Out My Editing Software
  • 10 Open Source Scrapers That Do What Paid APIs Do
  • Hermes Agent v0.20.0: It Finally Learned to Talk Back
  • PhotoGIMP: How I Turned GIMP into a Free Photoshop Clone
  • 8 Gemini Notebook Prompts That Actually Work
  • How I Built My Own Automation Hub (And the Problems That Nearly Stopped Me)
  • Hermes v0.19.1 Quietly Fixes the Frictions That Annoy You Most
  • Five AI Agents, One Trading Decision: The Architecture Behind the 95K Stars

Categories

  • DeepSeek
  • Qwen
  • GLM
  • Kimi‌
  • Codex
  • Hermes
  • Openclaw
  • Claude Code
  • Gemini
  • Hunyuan
  • China AI
  • AI Agent
  • AI Prompts
  • AI Tool Reviews
  • AI Guides
  • AI News

Tags

AI agent collaboration AI agent memory AI benchmarks AI coding assistant memory AI coding tools AI coding workflow AI context window AI dashboard AI deployment AI implementation AI models AI orchestration AI policy AI privacy AI security alternative AI hardware Anthropic ChatGPT Claude Claude coding Claude Tag Copilot cybersecurity developer tools FLUX GitHub code diagram knowledge management LLM LLM security local-first long context AI Midjourney Notion alternative Obsidian OpenAI OpenClaw open source open source AI persistent AI prompt-injection real AI coding agents Slack AI spreadsheet automation US government AI vetting workflow engine

About

Latest AI industry news and trend analysis, as well as tool evaluations.

Quick Links

  • About AIForker
  • Contact
  • How We Test
  • Privacy Policy
  • Tags

Category

  • AI NEWS
  • AI TOOL
  • AI GUIDES
  • CHINA AI
  • AI PROMPTS
Copyright2026 — AIForker.com. All rights reserved.