Low-cost local LLM inference: not just VRAM, but whole-machine data movement and scheduling
Local inference bottleneck has shifted from VRAM to whole-machine data movement + scheduling. Same card + model: llama.cpp ~50 tok/s vs Strata 70 tok/s (+40%). 3060 12GB whole-machine at ¥2500 = 77% performance, 38% price of 4090.
Source: araoai
Low-cost local LLM inference: not just VRAM, but whole-machine data movement and scheduling
This is a stance against the "VRAM-only" school — and the real dividing line for 2026 local inference.
Why "VRAM-only" is out
Old belief: fits in VRAM = runs well. Model 16GB + card 12GB = does not fit = cannot run.
In 2026 this changed — model weights can be offloaded to RAM/SSD. The question is no longer "can it fit" but "how fast can data move".
New bottleneck: whole-machine data movement + scheduling
Three data movement paths:
- GPU ↔ VRAM (memory bandwidth, HBM/GDDR)
- RTX 5090: 1.79 TB/s
- RTX 3060 12GB: 360 GB/s
- CPU ↔ RAM (DDR5 memory bandwidth)
- DDR5 5600 dual-channel: 89.6 GB/s
- DDR4 3200 dual-channel: 51.2 GB/s
- SSD ↔ CPU (NVMe sequential read)
- PCIe 4.0 NVMe: 7 GB/s
- PCIe 3.0 NVMe: 3.5 GB/s
- HDD: 0.2 GB/s (kills it)
Scheduling determines upper/lower bound:
| Framework | Scheduling | RTX 3060 12GB on 27B IQ2_XS |
|---|---|---|
| llama.cpp | Default sharding | ~50 tok/s |
| Strata | Three-tier offload auto-scheduling | 70 tok/s (+40%) |
| Strata + MTP | Speculative decoding | 79 tok/s (+58%) |
Scheduling gap = 40-58% — cheaper than "adding a card".
Whole-machine vs single card
| Setup | Capability | 27B tok/s | Cost |
|---|---|---|---|
| Single 24GB card (4090) | GPU only | 91 tok/s | ¥12000 |
| 12GB (3060) + 64GB DDR5 + NVMe | Three-tier offload | 70 tok/s | ¥4500 |
| 6GB (used 1080) + 64GB DDR5 + NVMe | Three-tier offload (forward-looking) | Close to 5090 (Strata v0.6) | ¥1500 |
3060 12GB whole-machine = 70 tok/s, 38% of 4090 cost, 77% of performance.
Upgrade path (existing PC)
By cost-benefit ratio:
| Priority | Upgrade | Cost | Gain |
|---|---|---|---|
| 🥇 | RAM: DDR4 16G → DDR5 32G | ¥600 | +30% tok/s |
| 🥈 | SSD: HDD → NVMe 4.0 | ¥500 | +20% tok/s |
| 🥉 | GPU: 8GB → 12GB | ¥1500 | +15% tok/s |
| 4️⃣ | GPU: 12GB → 24GB | ¥10000 | +30% tok/s |
Best cost-benefit: add RAM + swap SSD first, do not change card.
Evidence
- 91/70/79 tok/s: local DB benchmark_record id 51, 52, 53
- VRAM bandwidth: NVIDIA official spec
- DDR5/DDR4 bandwidth: JEDEC official spec
- Used 1080 close to 5090: Niko Veit @coldniko 2026-10-06 tweet (forward-looking)
Sources
- Whole-machine stance: combined Strata measurements + hardware specs
- Upgrade priority: based on tok/s delta vs upgrade cost ratio