[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f1ncpcnvtc5ccz":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":12,"bodyZh":13,"bodyEn":14,"sourceName":15,"sourceUrl":15,"updatedAt":16},"llamacpp-manual","GUIDE","2026-09-20T00:00:00.000Z","llama.cpp 手动部署","Manual setup with llama.cpp","下载 GGUF、自选量化与参数，理解每一步在做什么——可控性与性能天花板都更高。","Download GGUF weights, pick your own quantization and flags, and understand every step — more control, higher performance ceiling.",[],[],"## 1. 拿到 llama.cpp\n\n从 GitHub Releases 下载对应平台的预编译包，里面含 llama-server、llama-bench 等全套工具；需要特定 GPU 后端或最新特性时，也可以按官方文档源码编译。\n\n```\n# https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Freleases\n```\n\n## 2. 下载 GGUF 模型，选量化档位\n\nHugging Face 搜「模型名 + GGUF」即可找到社区转换好的权重。档位怎么选：Q4_K_M 是通用均衡档（8B 约 4.7 GB），Q8_0 更接近原模型精度但体积接近翻倍；模型页通常附各档位体积与困惑度对照表。\n\n## 3. 起服务\n\nllama-server 启动一个带网页对话界面的本地服务：-ngl 把多少层放进 GPU（99 = 全部；显存不够就调小，如 24），-c 是上下文长度。浏览器打开 localhost:8080 即可对话。\n\n```\nllama-server -m Qwen3-8B-Q4_K_M.gguf -ngl 99 -c 8192\n# 对话界面：http:\u002F\u002Flocalhost:8080\n```\n\n## 4. 最影响速度的两个参数\n\nGPU 卸载层数（-ngl）决定算力落在哪：放不下显存的层会退回 CPU，速度断崖式下降。上下文长度（-c）决定 KV cache 占用：显存随长度线性上涨，拉太长会挤掉模型本身的驻留空间。从短上下文起步，逐步加到自己机器的边界。\n\n> 上下文对吞吐的挤压，公开实测常常不标注——本站「当前局限」里把它列为第一位，测自己的机器时值得单独验证这一项。\n\n## 5. 测出你机器的数字\n\nllama-bench 输出两组吞吐：pp 是提示词处理速度（对应本站的 Prefill），tg 是文本生成速度（对应本站的 Decode）。换量化档、改 -ngl、拉长 -c 各跑一遍，就能摸清自己机器的能力边界。\n\n```\nllama-bench -m Qwen3-8B-Q4_K_M.gguf -ngl 99\n```\n\n## 6. 把结果投给本站\n\n记下模型、硬件、框架版本、量化档位和 pp\u002Ftg 两组数字，通过「联系方式」页投给我们——带原始出处的实测经过人工核对后进入数据库，让下一个搜「4090 跑 32B」的人少走弯路。","## 1. Get llama.cpp\n\nDownload a prebuilt package for your platform from GitHub Releases — it includes llama-server, llama-bench, and the rest of the toolset. For a specific GPU backend or the latest features, build from source following the official docs.\n\n```\n# https:\u002F\u002Fgithub.com\u002Fggml-org\u002Fllama.cpp\u002Freleases\n```\n\n## 2. Download GGUF weights and pick a quant\n\nSearch \"model name + GGUF\" on Hugging Face for community-converted weights. Choosing a level: Q4_K_M is the general-purpose balance (8B ≈ 4.7 GB); Q8_0 stays closer to the original model at nearly double the size. Model pages usually include a per-quant size and quality table.\n\n## 3. Start the server\n\nllama-server starts a local server with a built-in web chat: -ngl sets how many layers go to the GPU (99 = all; lower it, e.g. 24, if VRAM runs out) and -c sets context length. Open localhost:8080 to chat.\n\n```\nllama-server -m Qwen3-8B-Q4_K_M.gguf -ngl 99 -c 8192\n# Web chat: http:\u002F\u002Flocalhost:8080\n```\n\n## 4. The two flags that matter most for speed\n\nGPU offload (-ngl) decides where compute happens: layers that do not fit fall back to the CPU and speed drops off a cliff. Context length (-c) decides KV cache usage: VRAM grows linearly with length, and pushing it too far squeezes out the model itself. Start short and walk up to your machine’s limit.\n\n> Public benchmarks often omit the context length they tested at — this site lists it first among \"known limitations\". Worth verifying explicitly on your own machine.\n\n## 5. Benchmark your machine\n\nllama-bench reports two throughputs: pp (prompt processing — this site’s Prefill) and tg (text generation — this site’s Decode). Run it per quant level, per -ngl, per -c setting and you have your machine’s real envelope.\n\n```\nllama-bench -m Qwen3-8B-Q4_K_M.gguf -ngl 99\n```\n\n## 6. Send the results back\n\nNote the model, hardware, framework version, quantization, and your pp\u002Ftg numbers, then submit them via the Contact page — verified public benchmarks with cited sources enter the database, so the next person searching \"32B on a 4090\" starts from data.",null,"2026-10-01T04:48:19.343Z"]