← News
Guides

Low-cost local LLM inference: not just VRAM, but whole-machine data movement and scheduling

Local inference bottleneck has shifted from VRAM to whole-machine data movement + scheduling. Same card + model: llama.cpp ~50 tok/s vs Strata 70 tok/s (+40%). 3060 12GB whole-machine at ¥2500 = 77% performance, 38% price of 4090.

Source: araoai

Low-cost local LLM inference: not just VRAM, but whole-machine data movement and scheduling

This is a stance against the "VRAM-only" school — and the real dividing line for 2026 local inference.

Why "VRAM-only" is out

Old belief: fits in VRAM = runs well. Model 16GB + card 12GB = does not fit = cannot run.

In 2026 this changed — model weights can be offloaded to RAM/SSD. The question is no longer "can it fit" but "how fast can data move".

New bottleneck: whole-machine data movement + scheduling

Three data movement paths:

  1. GPU ↔ VRAM (memory bandwidth, HBM/GDDR)
    • RTX 5090: 1.79 TB/s
    • RTX 3060 12GB: 360 GB/s
  2. CPU ↔ RAM (DDR5 memory bandwidth)
    • DDR5 5600 dual-channel: 89.6 GB/s
    • DDR4 3200 dual-channel: 51.2 GB/s
  3. SSD ↔ CPU (NVMe sequential read)
    • PCIe 4.0 NVMe: 7 GB/s
    • PCIe 3.0 NVMe: 3.5 GB/s
    • HDD: 0.2 GB/s (kills it)

Scheduling determines upper/lower bound:

Framework Scheduling RTX 3060 12GB on 27B IQ2_XS
llama.cpp Default sharding ~50 tok/s
Strata Three-tier offload auto-scheduling 70 tok/s (+40%)
Strata + MTP Speculative decoding 79 tok/s (+58%)

Scheduling gap = 40-58% — cheaper than "adding a card".

Whole-machine vs single card

Setup Capability 27B tok/s Cost
Single 24GB card (4090) GPU only 91 tok/s ¥12000
12GB (3060) + 64GB DDR5 + NVMe Three-tier offload 70 tok/s ¥4500
6GB (used 1080) + 64GB DDR5 + NVMe Three-tier offload (forward-looking) Close to 5090 (Strata v0.6) ¥1500

3060 12GB whole-machine = 70 tok/s, 38% of 4090 cost, 77% of performance.

Upgrade path (existing PC)

By cost-benefit ratio:

Priority Upgrade Cost Gain
🥇 RAM: DDR4 16G → DDR5 32G ¥600 +30% tok/s
🥈 SSD: HDD → NVMe 4.0 ¥500 +20% tok/s
🥉 GPU: 8GB → 12GB ¥1500 +15% tok/s
4️⃣ GPU: 12GB → 24GB ¥10000 +30% tok/s

Best cost-benefit: add RAM + swap SSD first, do not change card.

Evidence

  • 91/70/79 tok/s: local DB benchmark_record id 51, 52, 53
  • VRAM bandwidth: NVIDIA official spec
  • DDR5/DDR4 bandwidth: JEDEC official spec
  • Used 1080 close to 5090: Niko Veit @coldniko 2026-10-06 tweet (forward-looking)

Sources

  • Whole-machine stance: combined Strata measurements + hardware specs
  • Upgrade priority: based on tok/s delta vs upgrade cost ratio