0
Quantized builds
36
Benchmark records
Quantized builds
Community and official quantized releases; "linked records" counts only benchmarks tied to that exact release.
No quantized builds cataloged yet.
Benchmark data
36 configurations
| Hardware | Framework | Quant | Decode | VRAM used | Samples | Evidence |
|---|---|---|---|---|---|---|
| NVIDIA RTX 5080 SUPER | vLLM | IQ2_XS | 90.7 tok/s | 11 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5080 SUPER | vLLM | Q5_K_M | 41.4 tok/s | 23.5 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5080 SUPER | vLLM | Q4_K_M | 50.6 tok/s | 19.3 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5080 SUPER | llama.cpp | IQ2_XS | 83.2 tok/s | 11 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5080 SUPER | llama.cpp | Q5_K_M | 38 tok/s | 23.5 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5080 SUPER | llama.cpp | Q4_K_M | 46.5 tok/s | 19.3 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | vLLM | IQ2_XS | 72.5 tok/s | 11 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | vLLM | Q5_K_M | 33.1 tok/s | 18 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | vLLM | Q4_K_M | 40.5 tok/s | 18 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | llama.cpp | IQ2_XS | 66.6 tok/s | 11 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | llama.cpp | Q5_K_M | 30.4 tok/s | 18 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | llama.cpp | Q4_K_M | 37.2 tok/s | 18 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4080 Super | vLLM | IQ2_XS | 59.6 tok/s | 11 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4080 Super | vLLM | Q5_K_M | 27.2 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4080 Super | vLLM | Q4_K_M | 33.3 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4080 Super | llama.cpp | IQ2_XS | 54.7 tok/s | 11 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4080 Super | llama.cpp | Q5_K_M | 25 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4080 Super | llama.cpp | Q4_K_M | 30.5 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 Ti | vLLM | IQ2_XS | 72.5 tok/s | 11 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 Ti | vLLM | Q5_K_M | 33.1 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 Ti | vLLM | Q4_K_M | 40.5 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 Ti | llama.cpp | IQ2_XS | 66.6 tok/s | 11 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 Ti | llama.cpp | Q5_K_M | 30.4 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 Ti | llama.cpp | Q4_K_M | 37.2 tok/s | 16 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | vLLM | IQ2_XS | 145.1 tok/s | 11 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | vLLM | Q5_K_M | 66.2 tok/s | 23.5 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | vLLM | Q4_K_M | 81 tok/s | 19.3 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | llama.cpp | IQ2_XS | 133.1 tok/s | 11 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | llama.cpp | Q5_K_M | 60.8 tok/s | 23.5 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | llama.cpp | Q4_K_M | 74.3 tok/s | 19.3 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | vLLM | IQ2_XS | 81.6 tok/s | 11 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | vLLM | Q5_K_M | 37.3 tok/s | 23.5 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | vLLM | Q4_K_M | 45.6 tok/s | 19.3 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | llama.cpp | IQ2_XS | 74.9 tok/s | 11 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | llama.cpp | Q5_K_M | 34.2 tok/s | 23.5 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | llama.cpp | Q4_K_M | 41.8 tok/s | 19.3 GB | 1 | L0 Self-reported |
36 of these records are not linked to a specific quantized build; the full set is on the search page.
Community discussion
No discussions yet — start the first one.
Start a discussion