Curated internet, served daily

A 27B Multimodal Model Squeezed Into 5.9 Gigabytes

PrismML's ternary 27B multimodal model lands at 5.9GB with 98.2% of full-precision benchmark performance, and the retention number is the part worth checking.

The compression is real, and the caveat sits in the benchmark math. PrismML’s Ternary Bonsai 2 27B stores weights as ternary values — −1, 0, +1 — with FP16 group-wise scaling, giving 1.76 effective bits per weight and a total footprint of 5.9GB. Nothing in that line is vague enough to hide behind.

It is built on Qwen3.8 27B and keeps the deployment profile of the first Bonsai release from two months ago: a 27B-class multimodal model that runs on a local device. The low-bit representation is applied end to end, and it carries a 262K-token context window, text-and-image input, and an Apache 2.0 license.

Against its full-precision counterpart it is more than 9x smaller while retaining 98.2% of aggregate benchmark performance at a suite score of 83.9. That retention figure is the fine print: reasoning, math, coding, instruction following, vision, and agentic tool use all collapse into one number, and the single task a deployment exists for can slip while the aggregate holds.

My own test would be that 262K context under ternary weights, since compression tends to survive short prompts and show wear over long ones. The picture I am waiting for is the 5.9GB file on a laptop that has not been upgraded in three years, loading without fuss. Run the long-context case first.

model-compression local-ai multimodal benchmarks

← Back to Daily