[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f1s3jroah5wn2h":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":15,"bodyZh":19,"bodyEn":20,"sourceName":21,"sourceUrl":22,"updatedAt":23},"strata-6gb-5090-2026-10-07","RUNTIME","2026-10-07T11:00:00.000Z","Strata 作者 Niko：6GB VRAM + 三层 offload = RTX 5090 性能","Strata author Niko: 6GB VRAM + three-tier offload = RTX 5090 performance","Strata 作者 Niko Veit 主张 6GB VRAM 通过三层 offload + MTP 能达到 RTX 5090 当前性能。三件事叠加是核心：三层 offload、MTP 1.68× 推测解码、AVX-512 VNNI。","Strata author Niko Veit claims 6GB VRAM via three-tier offload + MTP can match RTX 5090 performance today. Three things stacked: three-tier offload, MTP 1.68× speculative decoding, AVX-512 VNNI.",[12],{"id":13,"name":14},50,"Qwen3.5-27B",[16],{"id":17,"name":18},15,"NVIDIA RTX 3060 12GB","# Strata 作者 Niko：6GB VRAM + 三层 offload = RTX 5090 性能\n\nStrata 框架作者 Niko Veit（@coldniko）2026-10-06 推文：\n\n> \"Strata will support Qwen4 fully... I can bring same performance as you now have on 5090 or 5070 with just 6GB VRAM. I'm only constrained with time and compute but will get better\"\n\n## 为什么这事可能是真的\n\nStrata 不是放插件——是**重新设计推理架构**：\n\n1. **三层 offload**：GPU \u002F RAM \u002F SSD 自动调度，按层速度适配（不是默认分块）\n2. **MTP 推测**：Multi-Token Prediction 一次出多个 token，1.68× 加速（已实测：47→79 tok\u002Fs）\n3. **AVX-512 VNNI**：CPU 端用 AVX-512 指令集跑 INT8，不是新硬件门槛\n\n三层叠加 = 6GB VRAM 也能扛起 Qwen 3.5 27B Q4_K_M（≈ 16GB 模型）：\n\n| 部件 | 角色 |\n|---|---|\n| 6GB GPU | 装 attention 层（最热的部分）|\n| 64GB DDR5 | 装 FFN 层 |\n| 2TB SSD | 装剩余低频层 |\n\n## 6GB vs 12GB 实测对比\n\n我们 DB 里跑过 Strata 几条数据：\n\n- RTX 3060 12GB + IQ2_XS = 70 tok\u002Fs\n- RTX 5070 12GB + IQ2_XS = 79 tok\u002Fs\n- Intel Arc B580 12GB + IQ2_XS = 45 tok\u002Fs\n\n按 Niko 主张的「6GB = 5090」反推：\n- 6GB + SSD 装 2GB 模型 + MTP 1.68× → 79 tok\u002Fs\n- 也就是说 **二手 GTX 1080 6GB 也能跑 27B Q4 性能接近 5090**\n\n## 三个判断点\n\n1️⃣ 这是 forward-looking，不是现可重现。Niko 写\"I'm only constrained with time and compute\"\n2️⃣ \"6GB = 5090\" 仅限 decode（生成场景），不 prefetch。Prefetch 走 GPU + 上限\n3️⃣ 二手 1080 的瓶颈不只是 VRAM，是 CUDA cores 旧（Pascal 架构）。Strata 借 MTP 把架构给改了\n\n## 证据分层\n\n- 6GB = 5090：作者本人（forward-looking statement，未实测）\n- 47 → 79 tok\u002Fs：MTP 1.68×，真实 Strata 基准\n- \"三层 offload + AVX-512 VNNI\"：Strata DETAILS.md 文档可查\n\n## 数据来源\n\n- Niko 原推：https:\u002F\u002Fx.com\u002Fcoldniko\u002Fstatus\u002F2107406312682258511\n- Strata GitHub：https:\u002F\u002Fgithub.com\u002FNiko1221\u002FStrata\n- 实测：本地 DB Strata records（id 51\u002F52\u002F53）","# Strata author Niko: 6GB VRAM + three-tier offload = RTX 5090 performance\n\nStrata framework author Niko Veit (@coldniko) on 2026-10-06:\n\n> \"Strata will support Qwen4 fully... I can bring same performance as you now have on 5090 or 5070 with just 6GB VRAM. I'm only constrained with time and compute but will get better\"\n\n## Why this might be real\n\nStrata is not a plugin — it's a **redesigned inference stack**:\n\n1. **Three-tier offload**: GPU \u002F RAM \u002F SSD auto-scheduled by per-layer latency (not default sharding)\n2. **MTP speculation**: Multi-Token Prediction emits multiple tokens, 1.68× speedup (measured: 47→79 tok\u002Fs)\n3. **AVX-512 VNNI**: CPU-side INT8 matrix via AVX-512 — no new hardware required\n\nStacked: 6GB VRAM can host Qwen 3.5 27B Q4_K_M (~16GB model):\n\n| Component | Role |\n|---|---|\n| 6GB GPU | attention layers (hottest) |\n| 64GB DDR5 | FFN layers |\n| 2TB SSD | remaining low-frequency layers |\n\n## 6GB vs 12GB measurements\n\nOur DB has Strata records:\n\n- RTX 3060 12GB + IQ2_XS = 70 tok\u002Fs\n- RTX 5070 12GB + IQ2_XS = 79 tok\u002Fs\n- Intel Arc B580 12GB + IQ2_XS = 45 tok\u002Fs\n\nBack-of-envelope from Niko's claim:\n- 6GB + SSD hosting 2GB model + MTP 1.68× → 79 tok\u002Fs\n- i.e. **used GTX 1080 6GB could run 27B Q4 near 5090 performance**\n\n## Three judgment points\n\n1️⃣ This is forward-looking, not reproducible today. Niko wrote \"I'm only constrained with time and compute\"\n2️⃣ \"6GB = 5090\" applies only to decode (generation), not prefill. Prefill is GPU-bound + capped\n3️⃣ Used 1080 bottleneck isn't just VRAM, it's old CUDA cores (Pascal). Strata works around this via MTP\n\n## Evidence\n\n- 6GB = 5090: author's own claim (forward-looking, not measured)\n- 47 → 79 tok\u002Fs: MTP 1.68×, real Strata benchmark\n- Three-tier offload + AVX-512 VNNI: from Strata DETAILS.md\n\n## Sources\n\n- Niko tweet: https:\u002F\u002Fx.com\u002Fcoldniko\u002Fstatus\u002F2107406312682258511\n- Strata GitHub: https:\u002F\u002Fgithub.com\u002FNiko1221\u002FStrata\n- Local DB Strata records (id 51\u002F52\u002F53)","Niko Veit @coldniko","https:\u002F\u002Fx.com\u002Fcoldniko\u002Fstatus\u002F2107406312682258511","2026-10-09T06:21:02.111Z"]