Qwen3.8-27B · NVFP4
This page aggregates 1 real-world runs of Qwen3.8-27B (NVFP4) on NVIDIA RTX PRO 6000 Blackwell with SGLang, contributed by 1 independent source platforms; metrics are averages of published measurements.
171.1 tok/s
Decode
Decode speed
10,877.7 tok/s
Prefill
Prefill speed
1.51 s
TTFT
Time to first token
88.32 GB
VRAM
VRAM usage
1
Measured runs
1
Independent sources
Other
Source platforms
5 days ago
Last verified
Performance
- SGLang · NVFP4 (current)Decode 171.1 · Prefill 10877.7 ·
Core figures
- Decode (avg)
- 171.1 tok/s
- Prefill (avg)
- 10,877.7 tok/s
- TTFT (avg)
- 1.51 s
- VRAM (avg)
- 88.32 GB
- MTP acceptance rate
- —
- TTFB
- — GB
- Power draw
- — W
Configuration
Member-level fields are taken from the most recent run
- Model
- Qwen3.8-27B
- Quantization
- NVFP4
- Framework
- SGLang
- Version
- main (2026-09)
- Context length
- 262144 tokens
- Batch size
- 1
- Flash Attention
- On
Hardware
Nominal and measured figures are shown side by side; whether it runs is the reader's call
- GPU
- NVIDIA RTX PRO 6000 Blackwell
- Nominal VRAM
- 96 GB
- Measured VRAM (avg)
- 88.32 GB
- OS
- Ubuntu 24.04
Get started
Original model resources
Base model resources may not match this quantization. Use the verified deployment weights above when available.
Startup command
Command from the latest benchmark runsglang serve --model-path RadixArk/Qwen3.8-27B-NVFP4 --context-length 262144 --kv-cache-dtype fp8_e4m3 --mem-fraction-static 0.92 --attention-backend flashinfer --speculative-algorithm DFLASH --speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2 --speculative-num-draft-tokens 8Deployment guide
https://lcz.me/topic/1336Sources & evidence
1 measured records in total, each traceable to its original source
main (2026-09) · Ubuntu 24.04 · 262144 ctx
171.1 tok/s
Decode
10,877.7 tok/s
Prefill
1.51 s
TTFT
88.32 GB
VRAM
—
MTP
— W
Power draw
约 16K input、输出 4096、并发 1:TTFT 1.506s,prompt processing 10,877.7 t/s,decode 171.1 t/s。