NVIDIA RTX 5080

GPU

Blackwell 架构(GB203),16GB GDDR7 / 256-bit 位宽,显存带宽 960 GB/s,整卡功耗 360W。

16

运行记录

6

覆盖模型

4

覆盖框架

L2 跨框架验证 × 16

证据分布

典型性能表现

同硬件上各模型 × 框架 × 量化的已上架实测均值。

  1. llama.cpp · IQ4_XS Decode 133.14 · Prefill 731.5 ·
  2. Strata · Q3_K_M Decode 231.73 · Prefill 785.76 ·
  3. llama.cpp · IQ2_M Decode 338.41 · Prefill 1613.66 ·
  4. Strata · FP16 Decode 178.36 · Prefill 737.11 ·
  5. Strata · Q4_K_S Decode 363.26 · Prefill 1495.67 ·
  6. llama.cpp · IQ4_XS Decode 97.14 · Prefill 496.38 ·
  7. llama.cpp · Q3_K_M Decode 108.92 · Prefill 455.22 ·
  8. vLLM · IQ2_M Decode 131.31 · Prefill 559.09 ·
  9. llama.cpp · Q5_K_S Decode 91.75 · Prefill 454.66 ·
  10. Ollama · Q3_K_L Decode 101.88 · Prefill 405.42 ·
  11. Strata · IQ4_XS Decode 144.31 · Prefill 545.08 ·
  12. llama.cpp · Q4_K_S Decode 194.19 · Prefill 976.97 ·
  13. llama.cpp · Q8_0 Decode 121.42 · Prefill 409.38 ·
  14. Ollama · Q3_K_M Decode 198.41 · Prefill 879.72 ·
  15. Ollama · Q8_0 Decode 116.38 · Prefill 555.86 ·
  16. Ollama · IQ2_XS Decode 103.79 · Prefill 371.24 ·

核心规格

厂商
NVIDIA
架构
Blackwell
显存
16 GB
显存带宽
960 GB/s
TDP 功耗
360 W
发布日期
2025-01-30

运行记录

共 16 组配置

模型框架量化Decode生成耗时样本数
Qwen3.5-4Bllama.cppIQ4_XS133.1 tok/s—1
Qwen3.5-4BStrataQ3_K_M231.7 tok/s—1
Qwen3.5-0.8Bllama.cppIQ2_M338.4 tok/s—1
Qwen3.5-0.8BStrataFP16178.4 tok/s—1
Qwen3.5-0.8BStrataQ4_K_S363.3 tok/s—1
Qwen3.8-9Bllama.cppIQ4_XS97.1 tok/s—1
Qwen3.8-9Bllama.cppQ3_K_M108.9 tok/s—1
Qwen3.8-9BvLLMIQ2_M131.3 tok/s—1
Qwen3-8Bllama.cppQ5_K_S91.8 tok/s—1
Qwen3-8BOllamaQ3_K_L101.9 tok/s—1
Qwen3-8BStrataIQ4_XS144.3 tok/s—1
Qwen3-1.7Bllama.cppQ4_K_S194.2 tok/s—1
Qwen3-1.7Bllama.cppQ8_0121.4 tok/s—1
Qwen3-1.7BOllamaQ3_K_M198.4 tok/s—1
Qwen3-1.7BOllamaQ8_0116.4 tok/s—1
gemma-4-12B-itOllamaIQ2_XS103.8 tok/s—1