Strata hits 79 tok/s on 180B models with 12GB consumer GPUs: tiered memory + MTP speculative decoding
Strata's three-tier GPU/RAM/SSD memory architecture, MTP speculative decoding and AVX-512 VNNI kernels push Qwen3.8-Flash-Next (180B total / 6B active MoE) to 79 tok/s decode / 2090 tok/s prompt on RTX 5070 12GB + 64GB DDR5 at IQ2_XS.
What it is
Strata introduces a new paradigm: running 180B models at usable speeds on consumer 12GB GPUs. This is not simple CPU offload — it's a three-tier memory architecture:
- GPU/VRAM (12GB): caches hot experts + resident compute
- System RAM (64GB+): holds all 24,576 routed-expert copies (512 experts × 48 layers)
- SSD: stores the 28.8GB n-gram lookup table, read on demand via the OS page cache
Plus two critical optimizations:
- MTP speculative decoding: uses the model's built-in draft layer to verify multiple tokens at once, yielding a 1.6–1.8× speedup
- AVX-512 VNNI kernels: optimized for Q2_0 quantization, accelerating CPU-side weight unpacking
Measured numbers (RTX 5070 12GB + Ryzen 5 7600 + 64GB DDR5)
| Quant | Short decode | 128K-context decode | Prompt processing |
|---|---|---|---|
| Q2_0 | 93 tok/s | 74 tok/s | 2,170 tok/s |
| IQ2_XS (recommended) | 79 tok/s | 63 tok/s | 2,090 tok/s |
| IQ3_XXS | 62 tok/s | 49 tok/s | 1,750 tok/s |
| IQ3_S | 53 tok/s | 46 tok/s | 1,620 tok/s |
| Coder (IQ1_M) | 55 tok/s | 43 tok/s | 2,180 tok/s |
Why this is a real breakthrough
The conventional wisdom was "weights must all live in memory", which kept 180B models out of reach of 64GB-RAM workstations. Strata tiers weights down to SSD, pushing n-gram lookup tables (cold data) to disk while keeping the main compute path on RAM+GPU. That moves 180B on consumer hardware from "theoretically possible" to "daily usable".
Same-hardware comparison vs llama.cpp
| Framework | RTX 5070 12GB IQ2_XS decode |
|---|---|
| Strata | ~65.1 tok/s |
| llama.cpp | ~15 tok/s |
The gap comes from tiered memory + MTP speculation — two design choices llama.cpp doesn't have.
Local deployment angle
The minimum bar is a 12GB consumer GPU + 64GB DDR5 + an NVMe SSD; on the OS side, Linux + hugepages + IO scheduler set to none or mq-deadline to avoid page-cache thrashing on n-gram reads.