DeepSeek-V4-Flash-Vision-Exp · UD-Q8_K_XL
L0 Self-reported2x NVIDIA RTX 5090 · 64GB · llama.cpp
GitHub
Today
32.08 tok/s
Decode
671.43 tok/s
Prefill
0.12 s
TTFT
60.7 GB
VRAM
Config (latest run):
Ubuntu 24.04.4 LTS CUDA 13.3 UD-Q8_K_XL llama.cpp moecachev1.6-rc0
Sources: GitHub · 1 runs · 1 independent sources