0
Quantized builds
30
Benchmark records
Quantized builds
Community and official quantized releases; "linked records" counts only benchmarks tied to that exact release.
No quantized builds cataloged yet.
Benchmark data
30 configurations
| Hardware | Framework | Quant | Decode | VRAM used | Samples | Evidence |
|---|---|---|---|---|---|---|
| NVIDIA RTX 5080 SUPER | vLLM | Q8_0 | 29.4 tok/s | 24 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5080 SUPER | vLLM | Q5_K_M | 42.7 tok/s | 22.8 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5080 SUPER | vLLM | Q4_K_M | 52.3 tok/s | 18.7 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5080 SUPER | llama.cpp | Q8_0 | 27 tok/s | 24 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5080 SUPER | llama.cpp | Q5_K_M | 39.2 tok/s | 22.8 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5080 SUPER | llama.cpp | Q4_K_M | 48 tok/s | 18.7 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | vLLM | Q8_0 | 23.5 tok/s | 18 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | vLLM | Q5_K_M | 34.2 tok/s | 18 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | vLLM | Q4_K_M | 41.8 tok/s | 18 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | llama.cpp | Q8_0 | 21.6 tok/s | 18 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | llama.cpp | Q5_K_M | 31.3 tok/s | 18 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | llama.cpp | Q4_K_M | 38.4 tok/s | 18 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 Ti | vLLM | Q8_0 | 23.5 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 Ti | vLLM | Q5_K_M | 34.2 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 Ti | vLLM | Q4_K_M | 41.8 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 Ti | llama.cpp | Q8_0 | 21.6 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 Ti | llama.cpp | Q5_K_M | 31.3 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 Ti | llama.cpp | Q4_K_M | 38.4 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | vLLM | Q8_0 | 47 tok/s | 32 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | vLLM | Q5_K_M | 68.3 tok/s | 22.8 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | vLLM | Q4_K_M | 83.7 tok/s | 18.7 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | llama.cpp | Q8_0 | 43.1 tok/s | 32 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | llama.cpp | Q5_K_M | 62.7 tok/s | 22.8 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | llama.cpp | Q4_K_M | 76.8 tok/s | 18.7 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | vLLM | Q8_0 | 26.4 tok/s | 24 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | vLLM | Q5_K_M | 38.4 tok/s | 22.8 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | vLLM | Q4_K_M | 47.1 tok/s | 18.7 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | llama.cpp | Q8_0 | 24.3 tok/s | 24 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | llama.cpp | Q5_K_M | 35.3 tok/s | 22.8 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | llama.cpp | Q4_K_M | 43.2 tok/s | 18.7 GB | 1 | L0 Self-reported |
30 of these records are not linked to a specific quantized build; the full set is on the search page.
Community discussion
No discussions yet — start the first one.
Start a discussion