Qwen3.8-27B · UD-Q4_K_XL
This page aggregates 1 real-world runs of Qwen3.8-27B (UD-Q4_K_XL) on NVIDIA RTX PRO 4500 Blackwell with llama.cpp, contributed by 1 independent source platforms; metrics are averages of published measurements.
64.2 tok/s
Decode
Decode speed
1,330 tok/s
Prefill
Prefill speed
— s
TTFT
Time to first token
29 GB
VRAM
VRAM usage
1
Measured runs
1
Independent sources
Other
Source platforms
6 days ago
Last verified
Performance
- llama.cpp · UD-Q4_K_XL (current)Decode 64.2 · Prefill 1330 ·
Core figures
- Decode (avg)
- 64.2 tok/s
- Prefill (avg)
- 1,330 tok/s
- TTFT (avg)
- — s
- VRAM (avg)
- 29 GB
- MTP acceptance rate
- —
- TTFB
- — GB
- Power draw
- — W
Configuration
Member-level fields are taken from the most recent run
- Model
- Qwen3.8-27B
- Quantization
- UD-Q4_K_XL
- Framework
- llama.cpp
- Version
- Unsloth Desktop 0.1.801
- Context length
- 153856 tokens
- Batch size
- 1
- Flash Attention
- On
Hardware
Nominal and measured figures are shown side by side; whether it runs is the reader's call
- GPU
- NVIDIA RTX PRO 4500 Blackwell
- Nominal VRAM
- 32 GB
- Measured VRAM (avg)
- 29 GB
- OS
- Ubuntu
Get started
Weights for this deployment
Original model resources
Base model resources may not match this quantization. Use the verified deployment weights above when available.
Startup command
Command from the latest benchmark runllama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf --mmproj mmproj-F16.gguf --parallel 2 --flash-attn on -c 153856 -ngl -1 --kv-unified --spec-type draft-mtp --spec-draft-n-max 2Deployment guide
https://lcz.me/topic/1251Sources & evidence
1 measured records in total, each traceable to its original source
Unsloth Desktop 0.1.801 · Ubuntu · 153856 ctx
64.2 tok/s
Decode
1,330 tok/s
Prefill
— s
TTFT
29 GB
VRAM
—
MTP
— W
Power draw
RTX PRO4500,UD-Q4_K_XL,MTP n2、f16 KV。配置ctx153856、parallel2;实测slot0输入32798tok,预热后稳态prefill1330t/s、decode64.2t/s,原文f16内存约29GB。72.87%是全天195请求另组统计,非本测试,未填接受率。