Qwen3.6-27B · Q4_K_M
This page aggregates 1 real-world runs of Qwen3.6-27B (Q4_K_M) on NVIDIA RTX 3090 with llama.cpp, contributed by 1 independent source platforms; metrics are averages of published measurements.
57.02 tok/s
Decode
Decode speed
1,299.01 tok/s
Prefill
Prefill speed
— s
TTFT
Time to first token
23.44 GB
VRAM
VRAM usage
1
Measured runs
1
Independent sources
GitHub
Source platforms
5 days ago
Last verified
Performance
- llama.cpp · Q4_K_M (current)Decode 57.02 · Prefill 1299.01 ·
- llama.cpp · UD-Q3_K_XL Decode 36.67 · Prefill 39.44 ·
- llama.cpp · Q5_K_S Decode 46.57 · Prefill 1201.22 ·
Core figures
- Decode (avg)
- 57.02 tok/s
- Prefill (avg)
- 1,299.01 tok/s
- TTFT (avg)
- — s
- VRAM (avg)
- 23.44 GB
- MTP acceptance rate
- 0.65
- TTFB
- — GB
- Power draw
- 368 W
Configuration
Member-level fields are taken from the most recent run
- Model
- Qwen3.6-27B
- Quantization
- Q4_K_M
- Framework
- llama.cpp
- Version
- b9246
- Context length
- 262144 tokens
Hardware
Nominal and measured figures are shown side by side; whether it runs is the reader's call
- GPU
- NVIDIA RTX 3090
- Nominal VRAM
- 24 GB
- Measured VRAM (avg)
- 23.44 GB
- OS
- Ubuntu 24.04.4 LTS
- Driver
- 595.71.05
- CUDA
- 13.2
- Power draw
- 368 W
Get started
Original model resources
Base model resources may not match this quantization. Use the verified deployment weights above when available.
Sources & evidence
1 measured records in total, each traceable to its original source
b9246 · Ubuntu 24.04.4 LTS · CUDA 13.2 · 262144 ctx
57.02 tok/s
Decode
1,299.01 tok/s
Prefill
— s
TTFT
23.44 GB
VRAM
0.65
MTP
368 W
Power draw
0.48.418.321 I slot print_timing: id 0 | task 112 | prompt eval time = 111.50 ms / 60 tokens ( 1.86 ms per token, 538.13 tokens per second) 0.48.418.324 I slot print_timing: id 0 | task 112 | eval time = 35074.75 ms / 2000 tokens ( 17.54 ms per token, 57.02 tokens per second) 0.48.418.325 I slot print_timing: id 0 | task 112 | total time = 35186.25 ms / 2060 tokens