Runtime
Local inference trio: Strata / llama.cpp / vLLM decision framework
Pick by scenario, not by leaderboard: used GPU → Strata +40% speed, compatibility → llama.cpp GGUF, serving + concurrency → vLLM PagedAttention.
Local inference trio: Strata / llama.cpp / vLLM decision framework
After laying out what Strata is, Niko's 6GB claim, and used-GPU decisions, this one gives the install decision.
Three-way table
| Scenario | Pick | Why |
|---|---|---|
| Laptop / <8GB VRAM | llama.cpp | Mature, GGUF, docs |
| Used 1080/2060/3060 | Strata | Three-tier offload makes 6GB run 27B |
| Single 12-24GB card, 27B/30B | Strata | +40% speed (measured 70 vs 50 tok/s) |
| Multi-card 4090/5090 server | vLLM | PagedAttention + continuous batching |
| Compatibility first / wide model coverage | llama.cpp | GGUF de-facto standard |
| Serving / high-concurrency API | vLLM | Long context + mature scheduler |
| Want latest models (Qwen 4 / Llama 5) | llama.cpp | Strata v0.6 / vLLM not there yet |
| Experimenting / new ideas | Strata | v0.5.x is the research cut |
Why it's not "who is strongest"
Each solves a different problem:
- llama.cpp: broad compatibility — GGUF is the de-facto standard
- vLLM: throughput — PagedAttention reduces fragmentation, continuous batching
- Strata: small VRAM, big models — three-tier offload + MTP
Measurements (same model + card: Qwen 3.5 27B IQ2_XS)
| Framework | RTX 3060 12GB | RTX 5070 12GB | Arc B580 12GB |
|---|---|---|---|
| llama.cpp | ~50 tok/s | ~62 tok/s | ~32 tok/s |
| Strata | 70 tok/s | 79 tok/s | 45 tok/s |
| vLLM | won't fit (no native GGUF) | won't fit | won't fit |
Three judgment points
1️⃣ vLLM ↔ GGUF compatibility gap — vLLM eats HF safetensors. For GGUF, use llama.cpp 2️⃣ Strata v0.5.x model coverage is limited — Qwen 3.x / Llama 3.x / DeepSeek V3 work; Qwen 4 / Llama 5 full support = v0.6 3️⃣ All three stack: vLLM for serving + Strata for small models + llama.cpp for compatibility fallback