NVIDIA RTX 5080

GPU

NVIDIA Blackwell (GB203): 16 GB GDDR7 on a 256-bit bus, 960 GB/s of memory bandwidth, 360 W board power.

16

Run records

6

Models covered

4

Frameworks covered

L2 Cross-framework verified × 16

Evidence mix

Typical measured performance

Published benchmark averages per model × framework × quantization on this hardware.

  1. llama.cpp · IQ4_XS Decode 133.14 · Prefill 731.5 ·
  2. Strata · Q3_K_M Decode 231.73 · Prefill 785.76 ·
  3. llama.cpp · IQ2_M Decode 338.41 · Prefill 1613.66 ·
  4. Strata · FP16 Decode 178.36 · Prefill 737.11 ·
  5. Strata · Q4_K_S Decode 363.26 · Prefill 1495.67 ·
  6. llama.cpp · IQ4_XS Decode 97.14 · Prefill 496.38 ·
  7. llama.cpp · Q3_K_M Decode 108.92 · Prefill 455.22 ·
  8. vLLM · IQ2_M Decode 131.31 · Prefill 559.09 ·
  9. llama.cpp · Q5_K_S Decode 91.75 · Prefill 454.66 ·
  10. Ollama · Q3_K_L Decode 101.88 · Prefill 405.42 ·
  11. Strata · IQ4_XS Decode 144.31 · Prefill 545.08 ·
  12. llama.cpp · Q4_K_S Decode 194.19 · Prefill 976.97 ·
  13. llama.cpp · Q8_0 Decode 121.42 · Prefill 409.38 ·
  14. Ollama · Q3_K_M Decode 198.41 · Prefill 879.72 ·
  15. Ollama · Q8_0 Decode 116.38 · Prefill 555.86 ·
  16. Ollama · IQ2_XS Decode 103.79 · Prefill 371.24 ·

Key specs

Vendor
NVIDIA
Architecture
Blackwell
VRAM
16 GB
Memory bandwidth
960 GB/s
TDP
360 W
Release date
2025-01-30

Run records

16 configurations

ModelFrameworkQuantizationDecodeGen timeSamples
Qwen3.5-4Bllama.cppIQ4_XS133.1 tok/s—1
Qwen3.5-4BStrataQ3_K_M231.7 tok/s—1
Qwen3.5-0.8Bllama.cppIQ2_M338.4 tok/s—1
Qwen3.5-0.8BStrataFP16178.4 tok/s—1
Qwen3.5-0.8BStrataQ4_K_S363.3 tok/s—1
Qwen3.8-9Bllama.cppIQ4_XS97.1 tok/s—1
Qwen3.8-9Bllama.cppQ3_K_M108.9 tok/s—1
Qwen3.8-9BvLLMIQ2_M131.3 tok/s—1
Qwen3-8Bllama.cppQ5_K_S91.8 tok/s—1
Qwen3-8BOllamaQ3_K_L101.9 tok/s—1
Qwen3-8BStrataIQ4_XS144.3 tok/s—1
Qwen3-1.7Bllama.cppQ4_K_S194.2 tok/s—1
Qwen3-1.7Bllama.cppQ8_0121.4 tok/s—1
Qwen3-1.7BOllamaQ3_K_M198.4 tok/s—1
Qwen3-1.7BOllamaQ8_0116.4 tok/s—1
gemma-4-12B-itOllamaIQ2_XS103.8 tok/s—1