Qwen3.8-Flash-Next · UD-Q3_K_XL
L1 Reproduced4x NVIDIA Tesla P40 · 96GB · llama.cpp
GitHub
2 days ago
21.6 tok/s
Decode
330 tok/s
Prefill
— s
TTFT
— GB
VRAM
Config (latest run)
Linux UD-Q3_K_XL llama.cpp b11058
Sources: GitHub · 1 runs · 1 independent sources