← News
News

Same hardware, Strata 65 tok/s vs llama.cpp 15 tok/s — where does the gap come from?

On RTX 5070 12GB + 64GB DDR5, Strata hits ~65.1 tok/s vs llama.cpp's ~15 tok/s on Qwen3.8-Flash-Next IQ2_XS — the gap comes from SSD-streamed weights, MTP speculation, and AVX-512 VNNI kernels.

Source: Bittide AI Compass

What it is

On the same RTX 5070 12GB + Ryzen 5 7600 + 64GB DDR5-5600 workstation, the gap between Strata and llama.cpp on Qwen3.8-Flash-Next IQ2_XS is 4.3×:

Framework decode prompt
Strata ~65.1 tok/s ~2,090 tok/s
llama.cpp ~15 tok/s ~330 tok/s
gap 4.3× 6.3×

Three sources of the gap

  1. SSD-streamed weights: llama.cpp assumes weights must all live in RAM; with 180B and 64GB RAM it doesn't fit, causing heavy CPU↔RAM thrashing. Strata tiers weights to SSD (28.8GB n-gram + 64GB RAM for hot experts) and avoids that thrash.
  2. MTP speculative decoding: Qwen3.8-Flash-Next ships with a built-in draft layer; Strata consumes it directly for 1.6–1.8× speculation. llama.cpp's speculative decoding is still in an experimental branch, not on main.
  3. AVX-512 VNNI kernels: Strata has dedicated CPU unpacking paths for Q2_0 quantization, saturating modern desktop SIMD bandwidth. llama.cpp's quantization matrix kernels are more general-purpose; VNNI acceleration is under-used.

Why llama.cpp being slower ≠ "lagging"

llama.cpp's design goal is generality + cross-platform — runs on everything from 4-bit ARM CPUs to Apple Silicon to NVIDIA consumer cards. Strata specializes in Qwen3.8-Flash-Next + desktop CPU/GPU and pours all its optimization into that one target. The "specialist vs generalist" trade-off gets magnified to 4× in edge cases like 180B on 12GB.

Selection guidance

  • 180B on 12–16GB GPU + 64GB RAM: pick Strata
  • Llama / Mistral / DeepSeek and others: pick llama.cpp (widest ecosystem)
  • Multi-model workflow (5+ different models at once): pick llama.cpp / Ollama
  • Multi-GPU A100/H100 server-side: pick vLLM / SGLang

Local deployment angle

Strata only optimizes for Qwen3.8-Flash-Next (and future Qwen4 / qwen4exp-family) models. If you only use this generation's Qwen Flash tier, Strata is the strongest desktop option. If you also run other models, llama.cpp is the safer choice.