[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f3hhxxj1djyziu":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":21,"bodyZh":40,"bodyEn":41,"sourceName":42,"sourceUrl":43,"updatedAt":44},"data-movement-not-vram-2026-10-09","GUIDE","2026-10-09T14:00:00.000Z","大模型低成本运行：不只取决于显存，取决于整机数据搬运和调度能力","Low-cost local LLM inference: not just VRAM, but whole-machine data movement and scheduling","本地推理瓶颈已从 VRAM 转到整机数据搬运 + 调度。同卡同模型 llama.cpp ~50 tok\u002Fs vs Strata 70 tok\u002Fs（+40%）。3060 12GB 整机方案 ¥2500 = 4090 77% 性能 \u002F 38% 价格。","Local inference bottleneck has shifted from VRAM to whole-machine data movement + scheduling. Same card + model: llama.cpp ~50 tok\u002Fs vs Strata 70 tok\u002Fs (+40%). 3060 12GB whole-machine at ¥2500 = 77% performance, 38% price of 4090.",[12,15,18],{"id":13,"name":14},5,"Qwen3.8-27B",{"id":16,"name":17},6,"Qwen3.6-27B",{"id":19,"name":20},50,"Qwen3.5-27B",[22,25,28,31,34,37],{"id":23,"name":24},1,"NVIDIA RTX 4090",{"id":26,"name":27},13,"NVIDIA RTX 5090",{"id":29,"name":30},15,"NVIDIA RTX 3060 12GB",{"id":32,"name":33},18,"NVIDIA RTX 5070",{"id":35,"name":36},66,"NVIDIA GTX 1080 Ti",{"id":38,"name":39},94,"Intel Arc B580","# 大模型低成本运行：不只取决于显存，取决于整机数据搬运和调度能力\n\n这是一句反对\"显存党\"的判断——也是 2026 年本地推理的真正分水岭。\n\n## 为什么\"显存党\"out 了\n\n旧认知：能装下 = 跑得动。模型 16GB + 卡 12GB = 装不下 = 跑不动。\n\n2026 年变了——模型权重可以分块 offload 到 RAM\u002FSSD，问题不再是\"装得下\"，而是\"搬得快\"。\n\n## 新瓶颈：整机数据搬运 + 调度\n\n三个数据搬运路径：\n\n1. **GPU ↔ VRAM**（显存带宽，HBM\u002FGDDR 决定）\n   - RTX 5090: 1.79 TB\u002Fs\n   - RTX 3060 12GB: 360 GB\u002Fs\n2. **CPU ↔ RAM**（DDR5 内存带宽）\n   - DDR5 5600 双通道: 89.6 GB\u002Fs\n   - DDR4 3200 双通道: 51.2 GB\u002Fs\n3. **SSD ↔ CPU**（NVMe 顺序读）\n   - PCIe 4.0 NVMe: 7 GB\u002Fs\n   - PCIe 3.0 NVMe: 3.5 GB\u002Fs\n   - 机械硬盘: 0.2 GB\u002Fs（拖死）\n\n**调度决定上下限**：\n\n| 框架 | 调度方式 | RTX 3060 12GB 跑 27B IQ2_XS |\n|---|---|---|\n| llama.cpp | 默认分块 GPU+CPU | ~50 tok\u002Fs |\n| **Strata** | **三层 offload 自动调度** | **70 tok\u002Fs（+40%）** |\n| Strata + MTP | 推测解码 | 79 tok\u002Fs（+58%）|\n\n**调度差 = 40-58%**——比\"加一张卡\"便宜。\n\n## 整机 vs 单卡\n\n| 方案 | 整机能力 | 跑 27B tok\u002Fs | 整机成本 |\n|---|---|---|---|\n| 单卡 24GB（4090）| 仅 GPU | 91 tok\u002Fs | ¥12000 |\n| 单卡 12GB（3060）+ 64GB DDR5 + NVMe | 三层 offload | 70 tok\u002Fs | ¥4500 |\n| 单卡 6GB（1080 二手）+ 64GB DDR5 + NVMe | 三层 offload（forward-looking）| 接近 5090（待 Strata v0.6）| ¥1500 |\n\n3060 12GB 整机方案 = 70 tok\u002Fs，成本是 4090 的 38%，性能是 77%。\n\n## 升级路径（已有 PC）\n\n按\"成本-收益\"比排序：\n\n| 优先级 | 升级项 | 成本 | 收益 |\n|---|---|---|---|\n| 🥇 | RAM: DDR4 16G → DDR5 32G | ¥600 | +30% tok\u002Fs |\n| 🥈 | SSD: HDD → NVMe 4.0 | ¥500 | +20% tok\u002Fs |\n| 🥉 | GPU: 8GB → 12GB | ¥1500 | +15% tok\u002Fs |\n| 4️⃣ | GPU: 12GB → 24GB | ¥10000 | +30% tok\u002Fs |\n\n**性价比最高：先加 RAM + 换 SSD，不换卡**。\n\n## 证据分层\n\n- 91\u002F70\u002F79 tok\u002Fs：本地 DB benchmark_record id 51, 52, 53\n- 显存带宽：NVIDIA 官方 spec\n- DDR5\u002FDDR4 带宽：JEDEC 官方 spec\n- 二手 1080 接近 5090：Niko Veit @coldniko 2026-10-06 推文（forward-looking）\n\n## 数据来源\n\n- 整机党判断：综合 Strata 实测 + 硬件规格\n- 升级优先级：基于 tok\u002Fs 增量与升级成本比","# Low-cost local LLM inference: not just VRAM, but whole-machine data movement and scheduling\n\nThis is a stance against the \"VRAM-only\" school — and the real dividing line for 2026 local inference.\n\n## Why \"VRAM-only\" is out\n\nOld belief: fits in VRAM = runs well. Model 16GB + card 12GB = does not fit = cannot run.\n\nIn 2026 this changed — model weights can be offloaded to RAM\u002FSSD. The question is no longer \"can it fit\" but \"how fast can data move\".\n\n## New bottleneck: whole-machine data movement + scheduling\n\nThree data movement paths:\n\n1. **GPU ↔ VRAM** (memory bandwidth, HBM\u002FGDDR)\n   - RTX 5090: 1.79 TB\u002Fs\n   - RTX 3060 12GB: 360 GB\u002Fs\n2. **CPU ↔ RAM** (DDR5 memory bandwidth)\n   - DDR5 5600 dual-channel: 89.6 GB\u002Fs\n   - DDR4 3200 dual-channel: 51.2 GB\u002Fs\n3. **SSD ↔ CPU** (NVMe sequential read)\n   - PCIe 4.0 NVMe: 7 GB\u002Fs\n   - PCIe 3.0 NVMe: 3.5 GB\u002Fs\n   - HDD: 0.2 GB\u002Fs (kills it)\n\n**Scheduling determines upper\u002Flower bound**:\n\n| Framework | Scheduling | RTX 3060 12GB on 27B IQ2_XS |\n|---|---|---|\n| llama.cpp | Default sharding | ~50 tok\u002Fs |\n| **Strata** | **Three-tier offload auto-scheduling** | **70 tok\u002Fs (+40%)** |\n| Strata + MTP | Speculative decoding | 79 tok\u002Fs (+58%) |\n\n**Scheduling gap = 40-58%** — cheaper than \"adding a card\".\n\n## Whole-machine vs single card\n\n| Setup | Capability | 27B tok\u002Fs | Cost |\n|---|---|---|---|\n| Single 24GB card (4090) | GPU only | 91 tok\u002Fs | ¥12000 |\n| 12GB (3060) + 64GB DDR5 + NVMe | Three-tier offload | 70 tok\u002Fs | ¥4500 |\n| 6GB (used 1080) + 64GB DDR5 + NVMe | Three-tier offload (forward-looking) | Close to 5090 (Strata v0.6) | ¥1500 |\n\n3060 12GB whole-machine = 70 tok\u002Fs, 38% of 4090 cost, 77% of performance.\n\n## Upgrade path (existing PC)\n\nBy cost-benefit ratio:\n\n| Priority | Upgrade | Cost | Gain |\n|---|---|---|---|\n| 🥇 | RAM: DDR4 16G → DDR5 32G | ¥600 | +30% tok\u002Fs |\n| 🥈 | SSD: HDD → NVMe 4.0 | ¥500 | +20% tok\u002Fs |\n| 🥉 | GPU: 8GB → 12GB | ¥1500 | +15% tok\u002Fs |\n| 4️⃣ | GPU: 12GB → 24GB | ¥10000 | +30% tok\u002Fs |\n\n**Best cost-benefit: add RAM + swap SSD first, do not change card**.\n\n## Evidence\n\n- 91\u002F70\u002F79 tok\u002Fs: local DB benchmark_record id 51, 52, 53\n- VRAM bandwidth: NVIDIA official spec\n- DDR5\u002FDDR4 bandwidth: JEDEC official spec\n- Used 1080 close to 5090: Niko Veit @coldniko 2026-10-06 tweet (forward-looking)\n\n## Sources\n\n- Whole-machine stance: combined Strata measurements + hardware specs\n- Upgrade priority: based on tok\u002Fs delta vs upgrade cost ratio","araoai","https:\u002F\u002Faraoai.com\u002Fnews\u002Fdata-movement-not-vram-2026-10-09","2026-10-10T05:41:10.795Z"]