Qwen3-32B · Q4_K_M
This page aggregates 2 real-world runs of Qwen3-32B (Q4_K_M) on NVIDIA RTX 3090 with llama.cpp, contributed by 2 independent source platforms; metrics are averages of published measurements.
12.95 tok/s
Decode
Decode speed
143.35 tok/s
Prefill
Prefill speed
1.6 s
TTFT
Time to first token
19.5 GB
VRAM
VRAM usage
2
Measured runs
2
Independent sources
GitHub, Bilibili
Source platforms
2026-08-15
Last verified
Performance
- llama.cpp · Q4_K_M (current)Decode 12.95 · Prefill 143.35 ·
Core figures
- Decode (avg)
- 12.95 tok/s
- Prefill (avg)
- 143.35 tok/s
- TTFT (avg)
- 1.6 s
- VRAM (avg)
- 19.5 GB
- MTP acceptance rate
- —
- TTFB
- — GB
- Power draw
- 299.8 W
Configuration
Member-level fields are taken from the most recent run
- Model
- Qwen3-32B
- Quantization
- Q4_K_M
- Framework
- llama.cpp
- Version
- b4567
- Context length
- 8192 tokens
- GPU layers
- 99
Hardware
Nominal and measured figures are shown side by side; whether it runs is the reader's call
- GPU
- NVIDIA RTX 3090
- Nominal VRAM
- 24 GB
- Measured VRAM (avg)
- 19.5 GB
- OS
- Ubuntu 22.04
- Driver
- 535.183.01
- CUDA
- 12.2
- Power draw
- 298.4 W
Sources & evidence
2 measured records in total, each traceable to its original source
b4567 · Ubuntu 22.04 · CUDA 12.2 · 8192 ctx
12.8 tok/s
Decode
141.7 tok/s
Prefill
1.62 s
TTFT
19.5 GB
VRAM
—
MTP
298.4 W
Power draw
3090 上 Q4_K_M 约 13 tok/s。
b4602 · Ubuntu 22.04 · CUDA 12.2 · 8192 ctx
13.1 tok/s
Decode
145 tok/s
Prefill
1.58 s
TTFT
19.5 GB
VRAM
—
MTP
301.2 W
Power draw
3090 复现 13.1 tok/s。