Back to search results
Qwen3.8-27B · UD-Q8_K_XL
2x NVIDIA RTX 309048GBllama.cpp b10236L0 Self-reported
This page aggregates 1 real-world runs of Qwen3.8-27B (UD-Q8_K_XL) on 2x NVIDIA RTX 3090 with llama.cpp, contributed by 1 independent source platforms; metrics are averages of published measurements.
43.65 tok/s
Decode
Decode speed
1,600.18 tok/s
Prefill
Prefill speed
0.42 s
TTFT
Time to first token
44.12 GB
VRAM
VRAM usage
L0 Self-reported
1
Measured runs
1
Independent sources
GitHub
Source platforms
Today
Last verified
Performance
- llama.cpp · UD-Q8_K_XL (current)Decode 43.65 · Prefill 1600.18 ·
Core figures
- Decode (avg)
- 43.65 tok/s
- Prefill (avg)
- 1,600.18 tok/s
- TTFT (avg)
- 0.42 s
- VRAM (avg)
- 44.12 GB
- MTP acceptance rate
- 0.82
- TTFB
- — GB
- Power draw
- — W
Configuration
Member-level fields are taken from the most recent run
- Model
- Qwen3.8-27B
- Quantization
- UD-Q8_K_XL
- Framework
- llama.cpp
- Version
- b10236
- Context length
- 262144 tokens
- Batch size
- 512
- GPU layers
- 99
Hardware
Nominal and measured figures are shown side by side; whether it runs is the reader's call
- GPU
- 2x NVIDIA RTX 3090
- Nominal VRAM
- 48 GB
- Measured VRAM (avg)
- 44.12 GB
- OS
- Ubuntu 26.04 LTS
- Driver
- 610.57.04
- CUDA
- 13.3
Get started
Startup command
Command from the latest benchmark runmodel=/models/qwen3.8-27b/Qwen3.8-27B-UD-Q8_K_XL.gguf ubatch=512 kv=q8_0 ngl=99 split=0.55,0.45Sources & evidence
1 measured records in total, each traceable to its original source
b10236 · Ubuntu 26.04 LTS · CUDA 13.3 · 262144 ctx
43.65 tok/s
Decode
1,600.18 tok/s
Prefill
0.42 s
TTFT
44.12 GB
VRAM
0.82
MTP
— W
Power draw
[bench] narrative run 1/5: 1000 tok in 23.5s (43.22 decode TPS, ttft 394ms) [bench] narrative run 2/5: 1000 tok in 23.1s (44.20 decode TPS, ttft 516ms) [bench] narrative run 3/5: 1000 tok in 23.2s (43.81 decode TPS, ttft 395ms) [bench] narrative run 4/5: 991 tok in 23.3s (43.24 decode TPS, ttft 395ms) [bench] narrative run 5/5: 964 tok in 22.4s (43.77 decode TPS, ttft 392ms) [bench] prompt-processing run 1/1: 10004 prompt tok, 1600.2 PP tok/s, ttft 6252ms