Benchmark search

Try these searches

1 matching benchmark runs

ModelDeepSeek-V4-Flash-Vision-ExpHardware: AllFramework: AllQuantization: AllEvidence level: AllSource: All

DeepSeek-V4-Flash-Vision-Exp · UD-Q8_K_XL

L0 Self-reported

2x NVIDIA RTX 5090 · 64GB · llama.cpp

GitHub

Today

32.08 tok/s

Decode

671.43 tok/s

Prefill

0.12 s

TTFT

60.7 GB

VRAM

Config (latest run):

Ubuntu 24.04.4 LTS CUDA 13.3 UD-Q8_K_XL llama.cpp moecachev1.6-rc0

Sources: GitHub · 1 runs · 1 independent sources

1 results