27B en solo 5,9 GB: Bonsai 2 retiene el 98,2% y corre local (Apache 2.0)

27B en solo 5,9 GB: Bonsai 2 retiene el 98,2% y corre local (Apache 2.0)

21-09-2026 4:03:49
Compartir:

On September 17, 2026, PrismML (Pasadena) launched Ternary Bonsai 2 27B: a multimodal model with ~27.8 billion parameters compressed to just 5.9 GB — more than 9× less memory than full precision — retaining 98.2% of the base model’s aggregate score. Apache 2.0, free download starting today.

For SMBs and teams that want assistants, coding and agents without shipping everything to the cloud, that combo (27B + 5.9 GB + 98.2%) is the headline that matters: 27B-class quality with a consumer-GPU footprint.

The trick: 83.9 aggregate vs 85.4 base

Modelo con cabello cobrizo y top strap junto a dispositivo que ilustra un modelo 27B en 5,9 GB
Bonsai 2: 27B comprimido a 5,9 GB con 98,2% de retención

Per PR Newswire and PrismML docs, Bonsai 2 is based on Qwen3.8 27B. Across a 20-benchmark suite (reasoning, math, coding, instruction, vision, agentic tool use) it scores 83.9 vs 85.4 full-precision — 98.2% retention. The prior Ternary Bonsai 27B sat near ~95%.

Cited speed: up to ~143 tokens/sec on an NVIDIA GeForce RTX 5090. Trained on Google TPU v5.

What it is for (and what it is not)

Circuito y chip bajo luz cian: densidad de inteligencia en hardware de consumo
Más de 9× menos memoria vs precisión completa

PrismML targets local assistants, multimodal agents, coding, computer-use and long-horizon agentic workflows — cases where cloud hurts on latency, cost or privacy. This is not a toy 3B chatbot; it is a ternary 27B that fits in ~6 GB of model weights.

Dashboard de benchmarks agregados 83,9 vs 85,4 del modelo base
Suite de 20 benchmarks: razonamiento, coding, visión y tools

CEO Babak Hassibi frames an “intelligence density” thesis: not only bigger models in bigger datacenters, but useful capability per unit of memory, compute and energy. Advisor Ion Stoica (UC Berkeley) highlights how little is lost despite the footprint cut.

SMB checklist for local AI this week

Ingeniera probando IA local en laptop de oficina
Checklist pyme: RAG interno, coding y agentes sin nube
  1. Check whether your office GPU/CPU can load ~6 GB of weights + context.
  2. Pilot Bonsai 2 on a real flow: internal RAG, assisted coding or a local-tool agent.
  3. Compare monthly cloud spend vs power + already-amortized hardware.
  4. Review Apache 2.0 for commercial use without surprises.

If your SMB needs automation, a digital presence or an AI stack with data control, Presticorp builds tailored solutions. See presticorp.com/landingpages/pymes.

Writer’s take

Keep the “27B in 5.9 GB” headline and run a 48-hour pilot on a real business case. If 98.2% holds on your evals (not only PrismML’s suite), you just made private AI dramatically cheaper.

Sources

Compartir:

0 Comentarios

Deja un comentario

Landing pages especializadas

¿Proyecto totalmente personalizado? Contáctanos.

Si tu proyecto requiere una solución más enfocada, entra directo a la landing ideal para tu negocio y envíanos tu información en el formulario correspondiente.