[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$fimvzmk3cffek":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":12,"bodyZh":13,"bodyEn":14,"sourceName":15,"sourceUrl":15,"updatedAt":16},"ollama-quickstart","GUIDE","2026-09-20T00:00:00.000Z","Ollama 一键跑模型","Quickstart with Ollama","从安装到第一次对话的最快路径：一条命令把开源模型跑在自己机器上。","The fastest path from install to your first chat: run an open model on your own machine with one command.",[],[],"## 1. 安装 Ollama\n\nmacOS 与 Windows 直接从 ollama.com\u002Fdownload 下载安装包；Linux 一条命令完成安装。装完后在终端确认版本号，能打印即安装成功。\n\n```\ncurl -fsSL https:\u002F\u002Follama.com\u002Finstall.sh | sh\nollama -v\n```\n\n## 2. 跑第一个模型\n\n一条命令：首次运行会自动下载模型权重（8B 级别约 5 GB，视网速几分钟），完成后直接进入对话，输入 \u002Fbye 退出。\n这条命令背后发生了三件事：按量化档位下载 GGUF 权重、把模型加载进显存或内存、在本机启动推理服务——之后所有操作都围绕这套本地服务。\n\n```\nollama run qwen3:8b\n```\n\n## 3. 对照显存，选对规模\n\n显存不够时 Ollama 会把放不下的层挪进内存，能跑但明显变慢。Q4 量化、短上下文下的量级参考：4B 约 3 GB、8B 约 6 GB、14B 约 10 GB、32B 约 20 GB 显存。\n\n> 实际数字随上下文长度、并发与版本组合波动不小——这正是本站按记录标注实测环境、而不是只给一个均值的原因。\n\n## 4. 显存不够或没有独显\n\n先换更小的规模：4B 级别在集成显卡与纯 CPU 上也能跑。纯 CPU 生成速度大致降到独显的十几分之一，作答慢但完整可用，适合先体验再决定升级硬件。\n\n```\nollama run qwen3:4b\n```\n\n## 5. 当本地 API 用\n\nOllama 常驻后提供 OpenAI 兼容接口，把客户端或 IDE 插件的 base URL 指向本地地址即可接入。随时查看当前加载的模型与显存占用。\n\n```\nollama serve\n# OpenAI 兼容接口：http:\u002F\u002Flocalhost:11434\u002Fv1\nollama ps\n```\n\n## 6. 下一步\n\n想要更多控制——自定义上下文长度、挑量化档位、给机器测出基准数字——看第二篇《llama.cpp 手动部署》；跑出来的数字也欢迎通过「联系方式」投给本站，充实实测数据库。","## 1. Install Ollama\n\nmacOS and Windows: grab the installer from ollama.com\u002Fdownload. Linux: one command. Afterwards confirm the version prints in a terminal.\n\n```\ncurl -fsSL https:\u002F\u002Follama.com\u002Finstall.sh | sh\nollama -v\n```\n\n## 2. Run your first model\n\nOne command: the first run downloads the weights automatically (about 5 GB for an 8B model), then drops you straight into a chat. Type \u002Fbye to exit.\nBehind the scenes this does three things: downloads GGUF weights at a quantized size, loads the model into VRAM or system memory, and starts a local inference server — everything afterwards builds on that server.\n\n```\nollama run qwen3:8b\n```\n\n## 3. Match the model size to your VRAM\n\nWhen VRAM runs out, Ollama offloads layers to system memory — it still runs, but noticeably slower. Order-of-magnitude reference for Q4 quants with short contexts: 4B ≈ 3 GB, 8B ≈ 6 GB, 14B ≈ 10 GB, 32B ≈ 20 GB of VRAM.\n\n> Real numbers vary a lot with context length, concurrency, and version combinations — which is exactly why this site records the full test environment for every benchmark instead of publishing a bare average.\n\n## 4. No VRAM headroom, or no discrete GPU\n\nStep down a size: 4B-class models run on integrated graphics and on CPU alone. CPU-only generation is roughly an order of magnitude slower — slow but fully usable, fine for a first taste before deciding on hardware.\n\n```\nollama run qwen3:4b\n```\n\n## 5. Use it as a local API\n\nWhile running, Ollama serves an OpenAI-compatible endpoint: point a client or IDE extension at the local base URL and it just works. Check what is loaded and how much it occupies at any time.\n\n```\nollama serve\n# OpenAI-compatible endpoint: http:\u002F\u002Flocalhost:11434\u002Fv1\nollama ps\n```\n\n## 6. Next steps\n\nWant more control — custom context length, picking quantization levels, benchmarking your machine — continue with \"Manual setup with llama.cpp\". And if you end up with numbers of your own, send them in via the Contact page to grow the database.",null,"2026-10-01T04:48:19.327Z"]