Bonsai 27B is PrismML’s family of 1-bit and ternary Qwen3.6 27B models, with vision, reasoning, tool use, and a 262K-token context window.
To run the smallest Bonsai 27B, you need at least 4 GB of RAM.
Bonsai 27B models support tool use, vision input, and reasoning. They are available in gguf and mlx.
Bonsai 27B is PrismML's family of binary and ternary low-bit versions of Qwen3.6 27B. It retains the base model's vision, reasoning, and tool-use capabilities while making 27B-class local inference practical on much smaller devices.
| Variant | Weight representation | Deployed language-model footprint | Focus |
|---|---|---|---|
| 1-bit Bonsai 27B | Binary {−1, +1}, 1.125 effective bits per weight | ~3.9 GB | Smallest footprint |
| Ternary Bonsai 27B | Ternary {−1, 0, +1}, 1.71 effective bits per weight | ~7.2 GB | Higher quality retention |
PrismML applies the low-bit representation end to end across the language network, including embeddings, attention, MLPs, and the LM head. The headline footprints describe the language-model weights; actual runtime memory also includes the KV cache, activations, runtime buffers, and the optional vision tower, and grows with context length.
PrismML reports the following category averages from 15 benchmarks evaluated in thinking mode:
| Category | Qwen3.6 27B FP16 | Ternary Bonsai 27B | 1-bit Bonsai 27B |
|---|---|---|---|
| Math | 95.3 | 93.4 | 91.7 |
| Coding | 88.7 | 86.0 | 81.9 |
| Agentic and tool calling | 80.0 | 74.0 | 66.0 |
| Instruction following | 78.4 | 71.8 | 65.8 |
| Knowledge / STEM | 83.1 | 77.0 | 73.4 |
| Vision | 72.6 | 65.2 | 59.6 |
| Overall (15 benchmarks) | 85.0 | 80.5 | 76.1 |

Benchmark results and chart are reported by PrismML. See the Bonsai 27B announcement and whitepaper for the full per-benchmark results and methodology.
Bonsai 27B is available under the Apache 2.0 License.