Qwen3.8-27B · Q4_K_M
This page aggregates 1 real-world runs of Qwen3.8-27B (Q4_K_M) on AMD Radeon RX 7900 XTX with llama.cpp, contributed by 1 independent source platforms; metrics are averages of published measurements.
124.1 tok/s
Decode
Decode speed
156 tok/s
Prefill
Prefill speed
— s
TTFT
Time to first token
22.16 GB
VRAM
VRAM usage
1
Measured runs
1
Independent sources
Other
Source platforms
3 days ago
Last verified
Performance
- llama.cpp · Q4_K_M (current)Decode 124.1 · Prefill 156 ·
- llama.cpp · IQ3_S Decode 26.73 · Prefill 487.5 ·
Core figures
- Decode (avg)
- 124.1 tok/s
- Prefill (avg)
- 156 tok/s
- TTFT (avg)
- — s
- VRAM (avg)
- 22.16 GB
- MTP acceptance rate
- —
- TTFB
- — GB
- Power draw
- — W
Configuration
Member-level fields are taken from the most recent run
- Model
- Qwen3.8-27B
- Quantization
- Q4_K_M
- Framework
- llama.cpp
- Version
- 0.3.0-dev (ggml 0.22.0)
- Context length
- 204800 tokens
- Batch size
- 256
- GPU layers
- 99
- Flash Attention
- On
Hardware
Nominal and measured figures are shown side by side; whether it runs is the reader's call
- GPU
- AMD Radeon RX 7900 XTX
- Nominal VRAM
- 24 GB
- Measured VRAM (avg)
- 22.16 GB
- OS
- Ubuntu 24.04
- Driver
- Mesa/RADV 25.2.8
Get started
Original model resources
Base model resources may not match this quantization. Use the verified deployment weights above when available.
Startup command
Command from the latest benchmark run./build-vulkan/bin/llama-server -m models/Qwen3.8-27B-unsloth-Q4_K_M.gguf -c 204800 -fa on -np 1 -ub 256 --cache-type-k q5_1 --cache-type-v q4_0 --device Vulkan0 --spec-type ngram-map-k4v,draft-mtp --spec-draft-n-max 3 --spec-ngram-map-k4v-size-n 32 --jinjaSources & evidence
1 measured records in total, each traceable to its original source
0.3.0-dev (ggml 0.22.0) · Ubuntu 24.04 · 204800 ctx
124.1 tok/s
Decode
156 tok/s
Prefill
— s
TTFT
22.16 GB
VRAM
—
MTP
— W
Power draw
方案 B:長上下文(200K)把上面兩行 KV 改掉即可:-c 204800 --cache-type-k q5_1 --cache-type-v q4_0 → 124.1 / 68.1 / 40.9 t/s,VRAM 22.16 GiB