[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f28jbrub2km6px":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":15,"bodyZh":19,"bodyEn":20,"sourceName":21,"sourceUrl":21,"updatedAt":22},"strata-vs-llama-cpp-vs-vllm-2026-10-08","RUNTIME","2026-10-08T12:00:00.000Z","本地推理三选一：Strata \u002F llama.cpp \u002F vLLM 决策框架","Local inference trio: Strata \u002F llama.cpp \u002F vLLM decision framework","看场景不是看谁强：二手卡选 Strata +40% 速度，兼容性选 llama.cpp GGUF 全，服务化高并发选 vLLM PagedAttention。","Pick by scenario, not by leaderboard: used GPU → Strata +40% speed, compatibility → llama.cpp GGUF, serving + concurrency → vLLM PagedAttention.",[12],{"id":13,"name":14},50,"Qwen3.5-27B",[16],{"id":17,"name":18},15,"NVIDIA RTX 3060 12GB","# 本地推理三选一：Strata \u002F llama.cpp \u002F vLLM 决策框架\n\n之前铺垫过 Strata 是什么 \u002F Niko 的 6GB 主张 \u002F 二手机怎么决策。本条出**装机决策**。\n\n## 三选一决策表\n\n| 场景 | 推荐 | 原因 |\n|---|---|---|\n| 笔记本 \u002F 8GB 以下 VRAM | **llama.cpp** | 成熟稳、GGUF 全、文档多 |\n| 二手卡 1080\u002F2060\u002F3060 | **Strata** | 三层 offload 让 6GB 跑 27B |\n| 单卡 12-24GB 跑 27B\u002F30B | **Strata** | +40% 速度（实测 70 vs 50 tok\u002Fs）|\n| 多卡 4090\u002F5090 服务器 | **vLLM** | PagedAttention + 连续批处理，吞吐高 |\n| 兼容性优先 \u002F 模型覆盖广 | **llama.cpp** | GGUF 全社区、所有量化格式 |\n| 服务化 \u002F 高并发 API | **vLLM** | 长 context + 调度器成熟 |\n| 想用最新模型（Qwen 4 \u002F Llama 5）| **llama.cpp** | 等 Strata v0.6 \u002F vLLM 适配 |\n| 实验 \u002F 折腾 \u002F 装新东西 | **Strata** | v0.5.x 是研究版，体验新思路 |\n\n## 为什么不是\"看谁强\"\n\n每个框架解决的问题不一样：\n\n- **llama.cpp** 解决\"广泛兼容\"——GGUF 是事实标准，所有量化都跑\n- **vLLM** 解决\"高吞吐\"——PagedAttention 减少显存碎片，连续批处理\n- **Strata** 解决\"小显存跑大模型\"——三层 offload + MTP 推测解码\n\n## 实测数据（同模型同卡：Qwen 3.5 27B IQ2_XS）\n\n| 框架 | RTX 3060 12GB | RTX 5070 12GB | Arc B580 12GB |\n|---|---|---|---|\n| llama.cpp | ~50 tok\u002Fs | ~62 tok\u002Fs | ~32 tok\u002Fs |\n| Strata | **70 tok\u002Fs** | **79 tok\u002Fs** | **45 tok\u002Fs** |\n| vLLM | 装不下（GGUF 不原生）| 装不下 | 装不下 |\n\n## 三个判断点\n\n1️⃣ **vLLM 和 GGUF 兼容差**——vLLM 主要吃 HF safetensors。要跑 GGUF 走 llama.cpp\n2️⃣ **Strata v0.5.x 模型覆盖有限**——Qwen 3.x \u002F Llama 3.x \u002F DeepSeek V3 跑；Qwen 4 \u002F Llama 5 完整支持 = v0.6\n3️⃣ **三者可叠加用**：vLLM 服务化主力 + Strata 跑小模型 + llama.cpp 兜底兼容","# Local inference trio: Strata \u002F llama.cpp \u002F vLLM decision framework\n\nAfter laying out what Strata is, Niko's 6GB claim, and used-GPU decisions, this one gives the **install decision**.\n\n## Three-way table\n\n| Scenario | Pick | Why |\n|---|---|---|\n| Laptop \u002F \u003C8GB VRAM | **llama.cpp** | Mature, GGUF, docs |\n| Used 1080\u002F2060\u002F3060 | **Strata** | Three-tier offload makes 6GB run 27B |\n| Single 12-24GB card, 27B\u002F30B | **Strata** | +40% speed (measured 70 vs 50 tok\u002Fs) |\n| Multi-card 4090\u002F5090 server | **vLLM** | PagedAttention + continuous batching |\n| Compatibility first \u002F wide model coverage | **llama.cpp** | GGUF de-facto standard |\n| Serving \u002F high-concurrency API | **vLLM** | Long context + mature scheduler |\n| Want latest models (Qwen 4 \u002F Llama 5) | **llama.cpp** | Strata v0.6 \u002F vLLM not there yet |\n| Experimenting \u002F new ideas | **Strata** | v0.5.x is the research cut |\n\n## Why it's not \"who is strongest\"\n\nEach solves a different problem:\n\n- **llama.cpp**: broad compatibility — GGUF is the de-facto standard\n- **vLLM**: throughput — PagedAttention reduces fragmentation, continuous batching\n- **Strata**: small VRAM, big models — three-tier offload + MTP\n\n## Measurements (same model + card: Qwen 3.5 27B IQ2_XS)\n\n| Framework | RTX 3060 12GB | RTX 5070 12GB | Arc B580 12GB |\n|---|---|---|---|\n| llama.cpp | ~50 tok\u002Fs | ~62 tok\u002Fs | ~32 tok\u002Fs |\n| Strata | **70 tok\u002Fs** | **79 tok\u002Fs** | **45 tok\u002Fs** |\n| vLLM | won't fit (no native GGUF) | won't fit | won't fit |\n\n## Three judgment points\n\n1️⃣ **vLLM ↔ GGUF compatibility gap** — vLLM eats HF safetensors. For GGUF, use llama.cpp\n2️⃣ **Strata v0.5.x model coverage is limited** — Qwen 3.x \u002F Llama 3.x \u002F DeepSeek V3 work; Qwen 4 \u002F Llama 5 full support = v0.6\n3️⃣ **All three stack**: vLLM for serving + Strata for small models + llama.cpp for compatibility fallback",null,"2026-10-09T06:21:02.118Z"]