← News
Runtime

Local inference trio: Strata / llama.cpp / vLLM decision framework

Pick by scenario, not by leaderboard: used GPU → Strata +40% speed, compatibility → llama.cpp GGUF, serving + concurrency → vLLM PagedAttention.

Local inference trio: Strata / llama.cpp / vLLM decision framework

After laying out what Strata is, Niko's 6GB claim, and used-GPU decisions, this one gives the install decision.

Three-way table

Scenario Pick Why
Laptop / <8GB VRAM llama.cpp Mature, GGUF, docs
Used 1080/2060/3060 Strata Three-tier offload makes 6GB run 27B
Single 12-24GB card, 27B/30B Strata +40% speed (measured 70 vs 50 tok/s)
Multi-card 4090/5090 server vLLM PagedAttention + continuous batching
Compatibility first / wide model coverage llama.cpp GGUF de-facto standard
Serving / high-concurrency API vLLM Long context + mature scheduler
Want latest models (Qwen 4 / Llama 5) llama.cpp Strata v0.6 / vLLM not there yet
Experimenting / new ideas Strata v0.5.x is the research cut

Why it's not "who is strongest"

Each solves a different problem:

  • llama.cpp: broad compatibility — GGUF is the de-facto standard
  • vLLM: throughput — PagedAttention reduces fragmentation, continuous batching
  • Strata: small VRAM, big models — three-tier offload + MTP

Measurements (same model + card: Qwen 3.5 27B IQ2_XS)

Framework RTX 3060 12GB RTX 5070 12GB Arc B580 12GB
llama.cpp ~50 tok/s ~62 tok/s ~32 tok/s
Strata 70 tok/s 79 tok/s 45 tok/s
vLLM won't fit (no native GGUF) won't fit won't fit

Three judgment points

1️⃣ vLLM ↔ GGUF compatibility gap — vLLM eats HF safetensors. For GGUF, use llama.cpp 2️⃣ Strata v0.5.x model coverage is limited — Qwen 3.x / Llama 3.x / DeepSeek V3 work; Qwen 4 / Llama 5 full support = v0.6 3️⃣ All three stack: vLLM for serving + Strata for small models + llama.cpp for compatibility fallback