Qwen3.8-27B · IQ3_S
This page aggregates 2 real-world runs of Qwen3.8-27B (IQ3_S) on AMD Radeon RX 7900 XTX with llama.cpp, contributed by 1 independent source platforms; metrics are averages of published measurements.
26.73 tok/s
Decode
Decode speed
487.5 tok/s
Prefill
Prefill speed
— s
TTFT
Time to first token
22.15 GB
VRAM
VRAM usage
2
Measured runs
1
Independent sources
GitHub
Source platforms
5 days ago
Last verified
Performance
- llama.cpp · Q4_K_M Decode 124.1 · Prefill 156 ·
- llama.cpp · IQ3_S (current)Decode 26.73 · Prefill 487.5 ·
Core figures
- Decode (avg)
- 26.73 tok/s
- Prefill (avg)
- 487.5 tok/s
- TTFT (avg)
- — s
- VRAM (avg)
- 22.15 GB
- MTP acceptance rate
- —
- TTFB
- — GB
- Power draw
- — W
Configuration
Member-level fields are taken from the most recent run
- Model
- Qwen3.8-27B
- Quantization
- IQ3_S
- Framework
- llama.cpp
- Version
- b11160
- Context length
- 229376 tokens
- Flash Attention
- On
Hardware
Nominal and measured figures are shown side by side; whether it runs is the reader's call
- GPU
- AMD Radeon RX 7900 XTX
- Nominal VRAM
- 24 GB
- Measured VRAM (avg)
- 22.15 GB
- OS
- CachyOS (Arch)
Get started
Original model resources
Base model resources may not match this quantization. Use the verified deployment weights above when available.
Startup command
Command from the latest benchmark runllama-server -m <gguf> -c 229376 -ctk q8_0 -ctv q8_0 -fa on -np 1 --ctx-checkpoints 4 -ngl all --spec-type draft-mtp --spec-draft-n-max 2Sources & evidence
2 measured records in total, each traceable to its original source
b11160 · CachyOS (Arch) · 229376 ctx
22.03 tok/s
Decode
422 tok/s
Prefill
— s
TTFT
22.7 GB
VRAM
—
MTP
— W
Power draw
| 224K (`-c 229376`) | 221,167 | 422 | 22.3 / 25.1 / 18.7 | 22.7 GB | none (0 evicted) |
b11160 · CachyOS (Arch) · 204800 ctx
31.43 tok/s
Decode
553 tok/s
Prefill
— s
TTFT
21.6 GB
VRAM
—
MTP
— W
Power draw
| 200K (`-c 204800`, context only) | 127,983 | 553 | 30.6 / 36.0 / 27.7 | 21.6 GB | none |