27B en solo 5,9 GB: Bonsai 2 retiene el 98,2% y corre local (Apache 2.0)
On September 17, 2026, PrismML (Pasadena) launched Ternary Bonsai 2 27B: a multimodal model with ~27.8 billion parameters compressed to just 5.9 GB — more than 9× less memory than full precision — retaining 98.2% of the base model’s aggregate score. Apache 2.0, free download starting today.
For SMBs and teams that want assistants, coding and agents without shipping everything to the cloud, that combo (27B + 5.9 GB + 98.2%) is the headline that matters: 27B-class quality with a consumer-GPU footprint.
The trick: 83.9 aggregate vs 85.4 base

Per PR Newswire and PrismML docs, Bonsai 2 is based on Qwen3.8 27B. Across a 20-benchmark suite (reasoning, math, coding, instruction, vision, agentic tool use) it scores 83.9 vs 85.4 full-precision — 98.2% retention. The prior Ternary Bonsai 27B sat near ~95%.
Cited speed: up to ~143 tokens/sec on an NVIDIA GeForce RTX 5090. Trained on Google TPU v5.
What it is for (and what it is not)

PrismML targets local assistants, multimodal agents, coding, computer-use and long-horizon agentic workflows — cases where cloud hurts on latency, cost or privacy. This is not a toy 3B chatbot; it is a ternary 27B that fits in ~6 GB of model weights.

CEO Babak Hassibi frames an “intelligence density” thesis: not only bigger models in bigger datacenters, but useful capability per unit of memory, compute and energy. Advisor Ion Stoica (UC Berkeley) highlights how little is lost despite the footprint cut.
SMB checklist for local AI this week

- Check whether your office GPU/CPU can load ~6 GB of weights + context.
- Pilot Bonsai 2 on a real flow: internal RAG, assisted coding or a local-tool agent.
- Compare monthly cloud spend vs power + already-amortized hardware.
- Review Apache 2.0 for commercial use without surprises.
If your SMB needs automation, a digital presence or an AI stack with data control, Presticorp builds tailored solutions. See presticorp.com/landingpages/pymes.
Writer’s take
Keep the “27B in 5.9 GB” headline and run a 48-hour pilot on a real business case. If 98.2% holds on your evals (not only PrismML’s suite), you just made private AI dramatically cheaper.
Sources
Deja un comentario
Enviando comentario…
¿Proyecto totalmente personalizado? Contáctanos.
Si tu proyecto requiere una solución más enfocada, entra directo a la landing ideal para tu negocio y envíanos tu información en el formulario correspondiente.

0 Comentarios