Ternary Bonsai 2 27B

27BChatby Prism ML Start a discussion

A 27B fully-ternary quantisation from Prism ML, derived from Qwen3.8-27B with the architecture unchanged: language weights are ternary {−1, 0, +1} end to end (g128 groups with FP16 group-wise scaling, blockwise Hadamard rotation folded into the stored weights) at a real ~1.72 bits/weight — 5.95 GB as PTQ1_0 or 7.21 GB as PQ2_0, about 1/9.3 of FP16. Prism ML reports 98.2% of FP16 intelligence retained (84.78 average over 14 thinking-mode benchmarks), with the 262K context and ~75% linear-attention hybrid backbone inherited from the base model. Requires Prism ML's llama.cpp fork (CUDA / Metal / CPU); an MLX build also exists, and the vision tower is an optional Q8_0 mmproj pack. Apache-2.0.

0

Quantized builds

2

Benchmark records

1

Discussions

Quantized builds

Community and official quantized releases; "linked records" counts only benchmarks tied to that exact release.

No quantized builds cataloged yet.

Benchmark data

2 configurations

HardwareFrameworkQuantDecodeVRAM usedSamplesEvidence
NVIDIA RTX 5060 Ti 16GBllama.cppPTQ1_08.1 tok/s—1L1 Reproduced
NVIDIA RTX 3080llama.cppPTQ1_052.2 tok/s7.03 GB1L0 Self-reported

2 of these records are not linked to a specific quantized build; the full set is on the search page.

Community discussion

View all discussions