← News
News

Strata hits 79 tok/s on 180B models with 12GB consumer GPUs: tiered memory + MTP speculative decoding

Strata's three-tier GPU/RAM/SSD memory architecture, MTP speculative decoding and AVX-512 VNNI kernels push Qwen3.8-Flash-Next (180B total / 6B active MoE) to 79 tok/s decode / 2090 tok/s prompt on RTX 5070 12GB + 64GB DDR5 at IQ2_XS.

Source: Strata GitHub DETAILS.md + AGI Hunt

What it is

Strata introduces a new paradigm: running 180B models at usable speeds on consumer 12GB GPUs. This is not simple CPU offload — it's a three-tier memory architecture:

  • GPU/VRAM (12GB): caches hot experts + resident compute
  • System RAM (64GB+): holds all 24,576 routed-expert copies (512 experts × 48 layers)
  • SSD: stores the 28.8GB n-gram lookup table, read on demand via the OS page cache

Plus two critical optimizations:

  • MTP speculative decoding: uses the model's built-in draft layer to verify multiple tokens at once, yielding a 1.6–1.8× speedup
  • AVX-512 VNNI kernels: optimized for Q2_0 quantization, accelerating CPU-side weight unpacking

Measured numbers (RTX 5070 12GB + Ryzen 5 7600 + 64GB DDR5)

Quant Short decode 128K-context decode Prompt processing
Q2_0 93 tok/s 74 tok/s 2,170 tok/s
IQ2_XS (recommended) 79 tok/s 63 tok/s 2,090 tok/s
IQ3_XXS 62 tok/s 49 tok/s 1,750 tok/s
IQ3_S 53 tok/s 46 tok/s 1,620 tok/s
Coder (IQ1_M) 55 tok/s 43 tok/s 2,180 tok/s

Why this is a real breakthrough

The conventional wisdom was "weights must all live in memory", which kept 180B models out of reach of 64GB-RAM workstations. Strata tiers weights down to SSD, pushing n-gram lookup tables (cold data) to disk while keeping the main compute path on RAM+GPU. That moves 180B on consumer hardware from "theoretically possible" to "daily usable".

Same-hardware comparison vs llama.cpp

Framework RTX 5070 12GB IQ2_XS decode
Strata ~65.1 tok/s
llama.cpp ~15 tok/s

The gap comes from tiered memory + MTP speculation — two design choices llama.cpp doesn't have.

Local deployment angle

The minimum bar is a 12GB consumer GPU + 64GB DDR5 + an NVMe SSD; on the OS side, Linux + hugepages + IO scheduler set to none or mq-deadline to avoid page-cache thrashing on n-gram reads.