Back to search results

Qwen3.8-27B · UD-Q4_K_XL

Start a discussion
NVIDIA RTX PRO 4500 Blackwell32GBllama.cpp Unsloth Desktop 0.1.801L0 Self-reported

This page aggregates 1 real-world runs of Qwen3.8-27B (UD-Q4_K_XL) on NVIDIA RTX PRO 4500 Blackwell with llama.cpp, contributed by 1 independent source platforms; metrics are averages of published measurements.

64.2 tok/s

Decode

Decode speed

1,330 tok/s

Prefill

Prefill speed

— s

TTFT

Time to first token

29 GB

VRAM

VRAM usage

L0 Self-reported

1

Measured runs

1

Independent sources

Other

Source platforms

6 days ago

Last verified

Performance

  1. llama.cpp · UD-Q4_K_XL (current)Decode 64.2 · Prefill 1330 ·

Core figures

Decode (avg)
64.2 tok/s
Prefill (avg)
1,330 tok/s
TTFT (avg)
— s
VRAM (avg)
29 GB
MTP acceptance rate
—
TTFB
— GB
Power draw
— W

Configuration

Member-level fields are taken from the most recent run

Model
Qwen3.8-27B
Quantization
UD-Q4_K_XL
Framework
llama.cpp
Version
Unsloth Desktop 0.1.801
Context length
153856 tokens
Batch size
1
Flash Attention
On

Hardware

Nominal and measured figures are shown side by side; whether it runs is the reader's call

GPU
NVIDIA RTX PRO 4500 Blackwell
Nominal VRAM
32 GB
Measured VRAM (avg)
29 GB
OS
Ubuntu

Get started

Weights for this deployment

Original model resources

Base model resources may not match this quantization. Use the verified deployment weights above when available.

Startup command

Command from the latest benchmark run
llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf --mmproj mmproj-F16.gguf --parallel 2 --flash-attn on -c 153856 -ngl -1 --kv-unified --spec-type draft-mtp --spec-draft-n-max 2

Sources & evidence

1 measured records in total, each traceable to its original source

  1. L0 Self-reportedOtherOriginal link Verified on 2026-09-27

    Unsloth Desktop 0.1.801 · Ubuntu · 153856 ctx

    64.2 tok/s

    Decode

    1,330 tok/s

    Prefill

    — s

    TTFT

    29 GB

    VRAM

    —

    MTP

    — W

    Power draw

    RTX PRO4500,UD-Q4_K_XL,MTP n2、f16 KV。配置ctx153856、parallel2;实测slot0输入32798tok,预热后稳态prefill1330t/s、decode64.2t/s,原文f16内存约29GB。72.87%是全天195请求另组统计,非本测试,未填接受率。