← News
Guides

Strata install in 5 steps: 30 minutes to Qwen 3.5 27B

Five steps to install Strata and run Qwen 3.5 27B IQ2_XS at 70 tok/s. Rust toolchain → clone → release build → GGUF model → run + monitor.

Source: Strata GitHub

Strata install in 5 steps: 30 minutes to Qwen 3.5 27B

After covering what Strata is, used-GPU decisions, and the three-framework choice, this is the hands-on — 5 steps, 30 minutes to running.

5-step install

① Install Rust toolchain

Strata is Rust. Needs rustc + cargo.

curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
source "$HOME/.cargo/env"

⚠️ Rust ≥ 1.75. Below that cargo build will fail.

② Clone the Strata repo

git clone https://github.com/Niko1221/Strata.git
cd Strata
git checkout v0.5.x  # current stable

⚠️ Use the v0.5.x tag — main may have unreleased features.

③ Build (first time 5-10 min)

cargo build --release

⚠️ First build pulls hundreds of crates. Be patient. Don't cargo build without --release — debug mode runs 4× slower.

④ Download GGUF model

I use Qwen 3.5 27B IQ2_XS (~12GB quantized, 27B-class for general use).

huggingface-cli download Qwen/Qwen3.5-27B-Instruct-GGUF qwen3.5-27b-instruct.Q4_K_M.gguf --local-dir ./models
# or IQ2_XS for 12GB cards:
huggingface-cli download TheBloke/Qwen3.5-27B-Instruct-GGUF qwen3.5-27b-instruct.IQ2_XS.gguf --local-dir ./models

⚠️ IQ2_XS loses some quality but fits on 12GB. Q4_K_M needs 16GB+.

⑤ Run + monitor

strata --model ./models/qwen3.5-27b-instruct.IQ2_XS.gguf 
       --three-tier 
       --mtp-lookahead 4 
       --prompt "Write a Python hello world"

Add --monitor to open the layers-placement dashboard.

Measurements (RTX 3060 12GB + Qwen 3.5 27B IQ2_XS)

Stage tok/s
First start (5-15s) —
Prefill (first token) 1.2s
Decode stable 70 tok/s
Decode + MTP 79 tok/s (peak)
Three-tier offload running 30 tok/s (prefetch warmup)

Five judgment points

1️⃣ cargo build --release is mandatory — debug mode is 4× slower 2️⃣ 12GB VRAM needs IQ2_XS — Q4_K_M won't fit, OOM 3️⃣ SSD offload needs NVMe (>3 GB/s sequential read) — HDD will stall 4️⃣ First start is slow (5-15s) — normal model load + scheduler init 5️⃣ MTP 1.68× has limits — less effective on large KV cache, shorter prompts gain more

Evidence

  • Install steps: Strata GitHub README + DETAILS.md
  • 70 / 79 tok/s: local DB Strata record id 51
  • 5-15s start: Strata startup (model load + scheduler init)
  • 5-10 min build: Rust build on 3060 host + 8-core CPU

Sources

  • Strata GitHub: https://github.com/Niko1221/Strata
  • Local DB Strata record id 51
  • Rust: rustup.rs official installer
  • Qwen 3.5 27B GGUF: Hugging Face official repo