0
Quantized builds
18
Benchmark records
Quantized builds
Community and official quantized releases; "linked records" counts only benchmarks tied to that exact release.
No quantized builds cataloged yet.
Benchmark data
18 configurations
| Hardware | Framework | Quant | Decode | VRAM used | Samples | Evidence |
|---|---|---|---|---|---|---|
| NVIDIA RTX 5060 SUPER | llama.cpp | Q8_0 | 40.6 tok/s | 9.1 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5060 SUPER | llama.cpp | Q5_K_M | 59.2 tok/s | 6.4 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5060 SUPER | llama.cpp | Q4_K_M | 71.3 tok/s | 5.4 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | llama.cpp | Q8_0 | 81.3 tok/s | 9.1 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | llama.cpp | Q5_K_M | 118.5 tok/s | 6.4 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | llama.cpp | Q4_K_M | 142.6 tok/s | 5.4 GB | 1 | L0 Self-reported |
| NVIDIA RTX 3090 | llama.cpp | Q8_0 | 84.9 tok/s | 9.1 GB | 1 | L0 Self-reported |
| NVIDIA RTX 3090 | llama.cpp | Q5_K_M | 123.7 tok/s | 6.4 GB | 1 | L0 Self-reported |
| NVIDIA RTX 3090 | llama.cpp | Q4_K_M | 149 tok/s | 5.4 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 | llama.cpp | Q8_0 | 60.9 tok/s | 9.1 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 | llama.cpp | Q5_K_M | 88.8 tok/s | 6.4 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 | llama.cpp | Q4_K_M | 107 tok/s | 5.4 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | llama.cpp | Q8_0 | 162.5 tok/s | 9.1 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | llama.cpp | Q5_K_M | 236.9 tok/s | 6.4 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | llama.cpp | Q4_K_M | 285.3 tok/s | 5.4 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | llama.cpp | Q8_0 | 91.4 tok/s | 9.1 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | llama.cpp | Q5_K_M | 133.3 tok/s | 6.4 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | llama.cpp | Q4_K_M | 160.5 tok/s | 5.4 GB | 1 | L0 Self-reported |
18 of these records are not linked to a specific quantized build; the full set is on the search page.
Community discussion
No discussions yet — start the first one.
Start a discussion