Same hardware, Strata 65 tok/s vs llama.cpp 15 tok/s — where does the gap come from?
On RTX 5070 12GB + 64GB DDR5, Strata hits ~65.1 tok/s vs llama.cpp's ~15 tok/s on Qwen3.8-Flash-Next IQ2_XS — the gap comes from SSD-streamed weights, MTP speculation, and AVX-512 VNNI kernels.
Source: Bittide AI Compass
What it is
On the same RTX 5070 12GB + Ryzen 5 7600 + 64GB DDR5-5600 workstation, the gap between Strata and llama.cpp on Qwen3.8-Flash-Next IQ2_XS is 4.3×:
| Framework | decode | prompt |
|---|---|---|
| Strata | ~65.1 tok/s | ~2,090 tok/s |
| llama.cpp | ~15 tok/s | ~330 tok/s |
| gap | 4.3× | 6.3× |
Three sources of the gap
- SSD-streamed weights: llama.cpp assumes weights must all live in RAM; with 180B and 64GB RAM it doesn't fit, causing heavy CPU↔RAM thrashing. Strata tiers weights to SSD (28.8GB n-gram + 64GB RAM for hot experts) and avoids that thrash.
- MTP speculative decoding: Qwen3.8-Flash-Next ships with a built-in draft layer; Strata consumes it directly for 1.6–1.8× speculation. llama.cpp's speculative decoding is still in an experimental branch, not on main.
- AVX-512 VNNI kernels: Strata has dedicated CPU unpacking paths for Q2_0 quantization, saturating modern desktop SIMD bandwidth. llama.cpp's quantization matrix kernels are more general-purpose; VNNI acceleration is under-used.
Why llama.cpp being slower ≠ "lagging"
llama.cpp's design goal is generality + cross-platform — runs on everything from 4-bit ARM CPUs to Apple Silicon to NVIDIA consumer cards. Strata specializes in Qwen3.8-Flash-Next + desktop CPU/GPU and pours all its optimization into that one target. The "specialist vs generalist" trade-off gets magnified to 4× in edge cases like 180B on 12GB.
Selection guidance
- 180B on 12–16GB GPU + 64GB RAM: pick Strata
- Llama / Mistral / DeepSeek and others: pick llama.cpp (widest ecosystem)
- Multi-model workflow (5+ different models at once): pick llama.cpp / Ollama
- Multi-GPU A100/H100 server-side: pick vLLM / SGLang
Local deployment angle
Strata only optimizes for Qwen3.8-Flash-Next (and future Qwen4 / qwen4exp-family) models. If you only use this generation's Qwen Flash tier, Strata is the strongest desktop option. If you also run other models, llama.cpp is the safer choice.