2x NVIDIA RTX 5090
Desktop3
Run records
3
Models covered
1
Frameworks covered
L0 Self-reported × 3
Evidence mix
Typical measured performance
Published benchmark averages per model × framework × quantization on this hardware.
- llama.cpp · UD-Q8_K_XL Decode 32.08 · Prefill 671.43 ·
- llama.cpp · UD-IQ4_XS Decode 21.96 · Prefill 421.07 ·
- llama.cpp · UD-IQ2_XXS Decode 19.39 · Prefill 692.41 ·
Key specs
- VRAM
- 64 GB
Run records
3 configurations
| Model | Framework | Quantization | Decode | Samples |
|---|---|---|---|---|
| DeepSeek-V4-Flash-Vision-Exp | llama.cpp | UD-Q8_K_XL | 32.1 tok/s | 1 |
| GLM-5.3-Flash | llama.cpp | UD-IQ4_XS | 22 tok/s | 1 |
| DeepSeek-V4-Flash | llama.cpp | UD-IQ2_XXS | 19.4 tok/s | 1 |