Benchmark search

Try these searches

2 matching benchmark runs

ModelDeepSeek-V4-FlashHardware: AllFramework: AllQuantization: AllEvidence level: AllSource: All

DeepSeek-V4-Flash · UD-IQ2_XXS

L0 Self-reported

2x NVIDIA RTX 5090 · 64GB · llama.cpp

GitHub

Today

19.39 tok/s

Decode

692.41 tok/s

Prefill

0.12 s

TTFT

51.17 GB

VRAM

Config (latest run):

Ubuntu 24.04.4 LTS CUDA 13.3 UD-IQ2_XXS llama.cpp b10236

Sources: GitHub · 1 runs · 1 independent sources

DeepSeek-V4-Flash · UD-IQ2_XXS

L0 Self-reported

4x NVIDIA RTX 3090 · 96GB · llama.cpp

GitHub

Today

35.68 tok/s

Decode

420.73 tok/s

Prefill

0.11 s

TTFT

87.16 GB

VRAM

Config (latest run):

Ubuntu 24.04.4 LTS CUDA 13.2 UD-IQ2_XXS llama.cpp club3090-moecache-e0a2089

Sources: GitHub · 1 runs · 1 independent sources

2 results