[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f1u5qlfc02040y":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":15,"bodyZh":16,"bodyEn":17,"sourceName":18,"sourceUrl":19,"updatedAt":20},"strata-12gb-consumer-gpu-180b","NEWS","2026-10-05T04:12:00.000Z","Strata 让 180B 模型在 12GB 消费卡上跑出 79 tok\u002Fs：分层内存 + MTP 推测解码","Strata hits 79 tok\u002Fs on 180B models with 12GB consumer GPUs: tiered memory + MTP speculative decoding","Strata 通过 GPU\u002FRAM\u002FSSD 三层内存架构 + MTP 推测解码 + AVX-512 VNNI 内核，让 Qwen3.8-Flash-Next（180B 总 \u002F 6B 激活 MoE）在 RTX 5070 12GB + 64GB DDR5 上达到 IQ2_XS 79 tok\u002Fs decode \u002F 2090 tok\u002Fs prompt。","Strata's three-tier GPU\u002FRAM\u002FSSD memory architecture, MTP speculative decoding and AVX-512 VNNI kernels push Qwen3.8-Flash-Next (180B total \u002F 6B active MoE) to 79 tok\u002Fs decode \u002F 2090 tok\u002Fs prompt on RTX 5070 12GB + 64GB DDR5 at IQ2_XS.",[12],{"id":13,"name":14},11,"Qwen3.8-Flash-Next",[],"### 是什么\n\n[Strata](https:\u002F\u002Fgithub.com\u002FNiko1221\u002FStrata) 给出了一种新范式：**让 180B 模型在消费级 12GB 显卡上跑出可用的速度**。这不是简单的 CPU offload，而是三层内存协同：\n\n- **GPU\u002FVRAM（12GB）**：缓存最常用专家 + 常驻计算部分\n- **系统 RAM（64GB+）**：存放全部 24,576 个路由专家副本（512 专家 × 48 层）\n- **SSD**：存储 28.8GB 的 n-gram 查找表，由 OS page cache 按需读取\n\n叠加两项关键优化：\n- **MTP 推测解码**：用模型自带的 draft layer 一次验证多个 token，提升 1.6–1.8 倍\n- **AVX-512 VNNI 内核**：专为 Q2_0 量化优化，CPU 端权重解压加速\n\n### 实测数据（RTX 5070 12GB + Ryzen 5 7600 + 64GB DDR5）\n\n| 量化 | 短对话 decode | 128K 上下文 decode | prompt 处理 |\n|---|---|---|---|\n| Q2_0 | 93 tok\u002Fs | 74 tok\u002Fs | 2,170 tok\u002Fs |\n| **IQ2_XS**（推荐档） | **79 tok\u002Fs** | **63 tok\u002Fs** | **2,090 tok\u002Fs** |\n| IQ3_XXS | 62 tok\u002Fs | 49 tok\u002Fs | 1,750 tok\u002Fs |\n| IQ3_S | 53 tok\u002Fs | 46 tok\u002Fs | 1,620 tok\u002Fs |\n| Coder (IQ1_M) | 55 tok\u002Fs | 43 tok\u002Fs | 2,180 tok\u002Fs |\n\n### 为什么是真正的突破\n\n传统思路是\"权重必须全部在内存里\"，于是 180B 模型被挡在 64GB+ 内存工作站门外。Strata 把权重分层到 SSD，让 n-gram 查找表这种**冷数据**走磁盘，主干计算走 RAM+GPU。这把消费级硬件本地跑 180B 从\"理论可行\"推到\"日常可用\"。\n\n### 与 llama.cpp 同硬件对比\n\n| 框架 | RTX 5070 12GB IQ2_XS decode |\n|---|---|\n| Strata | ~65.1 tok\u002Fs |\n| llama.cpp | ~15 tok\u002Fs |\n\n差距来源主要是三层内存 + MTP 推测这两件 llama.cpp 没有的设计。\n\n### 本地部署角度\n\n最低门槛是 **12GB 消费卡 + 64GB DDR5 + 一块 NVMe SSD**；操作系统层面建议 Linux + 大页（hugepages）+ IO 调度器切到 `none` 或 `mq-deadline`，避免 page cache 抖动拖慢 n-gram 读取。","### What it is\n\n[Strata](https:\u002F\u002Fgithub.com\u002FNiko1221\u002FStrata) introduces a new paradigm: **running 180B models at usable speeds on consumer 12GB GPUs**. This is not simple CPU offload — it's a three-tier memory architecture:\n\n- **GPU\u002FVRAM (12GB)**: caches hot experts + resident compute\n- **System RAM (64GB+)**: holds all 24,576 routed-expert copies (512 experts × 48 layers)\n- **SSD**: stores the 28.8GB n-gram lookup table, read on demand via the OS page cache\n\nPlus two critical optimizations:\n- **MTP speculative decoding**: uses the model's built-in draft layer to verify multiple tokens at once, yielding a 1.6–1.8× speedup\n- **AVX-512 VNNI kernels**: optimized for Q2_0 quantization, accelerating CPU-side weight unpacking\n\n### Measured numbers (RTX 5070 12GB + Ryzen 5 7600 + 64GB DDR5)\n\n| Quant | Short decode | 128K-context decode | Prompt processing |\n|---|---|---|---|\n| Q2_0 | 93 tok\u002Fs | 74 tok\u002Fs | 2,170 tok\u002Fs |\n| **IQ2_XS** (recommended) | **79 tok\u002Fs** | **63 tok\u002Fs** | **2,090 tok\u002Fs** |\n| IQ3_XXS | 62 tok\u002Fs | 49 tok\u002Fs | 1,750 tok\u002Fs |\n| IQ3_S | 53 tok\u002Fs | 46 tok\u002Fs | 1,620 tok\u002Fs |\n| Coder (IQ1_M) | 55 tok\u002Fs | 43 tok\u002Fs | 2,180 tok\u002Fs |\n\n### Why this is a real breakthrough\n\nThe conventional wisdom was \"weights must all live in memory\", which kept 180B models out of reach of 64GB-RAM workstations. Strata tiers weights down to SSD, pushing n-gram lookup tables (cold data) to disk while keeping the main compute path on RAM+GPU. That moves 180B on consumer hardware from \"theoretically possible\" to \"daily usable\".\n\n### Same-hardware comparison vs llama.cpp\n\n| Framework | RTX 5070 12GB IQ2_XS decode |\n|---|---|\n| Strata | ~65.1 tok\u002Fs |\n| llama.cpp | ~15 tok\u002Fs |\n\nThe gap comes from tiered memory + MTP speculation — two design choices llama.cpp doesn't have.\n\n### Local deployment angle\n\nThe minimum bar is **a 12GB consumer GPU + 64GB DDR5 + an NVMe SSD**; on the OS side, Linux + hugepages + IO scheduler set to `none` or `mq-deadline` to avoid page-cache thrashing on n-gram reads.","Strata GitHub DETAILS.md + AGI Hunt","https:\u002F\u002Fagihunt.info\u002Fen\u002Fp\u002F1a0fd67e46da89223aadf6cd889","2026-10-05T02:23:19.764Z"]