A 27B fully-ternary quantisation from Prism ML, derived from Qwen3.8-27B with the architecture unchanged: language weights are ternary {−1, 0, +1} end to end (g128 groups with FP16 group-wise scaling, blockwise Hadamard rotation folded into the stored weights) at a real ~1.72 bits/weight — 5.95 GB as PTQ1_0 or 7.21 GB as PQ2_0, about 1/9.3 of FP16. Prism ML reports 98.2% of FP16 intelligence retained (84.78 average over 14 thinking-mode benchmarks), with the 262K context and ~75% linear-attention hybrid backbone inherited from the base model. Requires Prism ML's llama.cpp fork (CUDA / Metal / CPU); an MLX build also exists, and the vision tower is an optional Q8_0 mmproj pack. Apache-2.0.
0
Quantized builds
2
Benchmark records
1
Discussions
Quantized builds
Community and official quantized releases; "linked records" counts only benchmarks tied to that exact release.
No quantized builds cataloged yet.
Benchmark data
2 configurations
| Hardware | Framework | Quant | Decode | VRAM used | Samples | Evidence |
|---|---|---|---|---|---|---|
| NVIDIA RTX 5060 Ti 16GB | llama.cpp | PTQ1_0 | 8.1 tok/s | — | 1 | L1 Reproduced |
| NVIDIA RTX 3080 | llama.cpp | PTQ1_0 | 52.2 tok/s | 7.03 GB | 1 | L0 Self-reported |
2 of these records are not linked to a specific quantized build; the full set is on the search page.