2x NVIDIA RTX 5090

Desktop

3

Run records

3

Models covered

1

Frameworks covered

L0 Self-reported × 3

Evidence mix

Typical measured performance

Published benchmark averages per model × framework × quantization on this hardware.

  1. llama.cpp · UD-Q8_K_XL Decode 32.08 · Prefill 671.43 ·
  2. llama.cpp · UD-IQ4_XS Decode 21.96 · Prefill 421.07 ·
  3. llama.cpp · UD-IQ2_XXS Decode 19.39 · Prefill 692.41 ·

Key specs

VRAM
64 GB

Run records

3 configurations

ModelFrameworkQuantizationDecodeSamples
DeepSeek-V4-Flash-Vision-Expllama.cppUD-Q8_K_XL32.1 tok/s1
GLM-5.3-Flashllama.cppUD-IQ4_XS22 tok/s1
DeepSeek-V4-Flashllama.cppUD-IQ2_XXS19.4 tok/s1