Strata install in 5 steps: 30 minutes to Qwen 3.5 27B
Five steps to install Strata and run Qwen 3.5 27B IQ2_XS at 70 tok/s. Rust toolchain → clone → release build → GGUF model → run + monitor.
Source: Strata GitHub
Strata install in 5 steps: 30 minutes to Qwen 3.5 27B
After covering what Strata is, used-GPU decisions, and the three-framework choice, this is the hands-on — 5 steps, 30 minutes to running.
5-step install
① Install Rust toolchain
Strata is Rust. Needs rustc + cargo.
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
source "$HOME/.cargo/env"
⚠️ Rust ≥ 1.75. Below that cargo build will fail.
② Clone the Strata repo
git clone https://github.com/Niko1221/Strata.git
cd Strata
git checkout v0.5.x # current stable
⚠️ Use the v0.5.x tag — main may have unreleased features.
③ Build (first time 5-10 min)
cargo build --release
⚠️ First build pulls hundreds of crates. Be patient. Don't cargo build without --release — debug mode runs 4× slower.
④ Download GGUF model
I use Qwen 3.5 27B IQ2_XS (~12GB quantized, 27B-class for general use).
huggingface-cli download Qwen/Qwen3.5-27B-Instruct-GGUF qwen3.5-27b-instruct.Q4_K_M.gguf --local-dir ./models
# or IQ2_XS for 12GB cards:
huggingface-cli download TheBloke/Qwen3.5-27B-Instruct-GGUF qwen3.5-27b-instruct.IQ2_XS.gguf --local-dir ./models
⚠️ IQ2_XS loses some quality but fits on 12GB. Q4_K_M needs 16GB+.
⑤ Run + monitor
strata --model ./models/qwen3.5-27b-instruct.IQ2_XS.gguf
--three-tier
--mtp-lookahead 4
--prompt "Write a Python hello world"
Add --monitor to open the layers-placement dashboard.
Measurements (RTX 3060 12GB + Qwen 3.5 27B IQ2_XS)
| Stage | tok/s |
|---|---|
| First start (5-15s) | — |
| Prefill (first token) | 1.2s |
| Decode stable | 70 tok/s |
| Decode + MTP | 79 tok/s (peak) |
| Three-tier offload running | 30 tok/s (prefetch warmup) |
Five judgment points
1️⃣ cargo build --release is mandatory — debug mode is 4× slower
2️⃣ 12GB VRAM needs IQ2_XS — Q4_K_M won't fit, OOM
3️⃣ SSD offload needs NVMe (>3 GB/s sequential read) — HDD will stall
4️⃣ First start is slow (5-15s) — normal model load + scheduler init
5️⃣ MTP 1.68× has limits — less effective on large KV cache, shorter prompts gain more
Evidence
- Install steps: Strata GitHub README + DETAILS.md
- 70 / 79 tok/s: local DB Strata record id 51
- 5-15s start: Strata startup (model load + scheduler init)
- 5-10 min build: Rust build on 3060 host + 8-core CPU
Sources
- Strata GitHub: https://github.com/Niko1221/Strata
- Local DB Strata record id 51
- Rust: rustup.rs official installer
- Qwen 3.5 27B GGUF: Hugging Face official repo