Llama 5 series flagship; 405B parameters, matches or exceeds GPT-5 on most reasoning benchmarks.
0
Quantized builds
6
Benchmark records
Quantized builds
Community and official quantized releases; "linked records" counts only benchmarks tied to that exact release.
No quantized builds cataloged yet.
Benchmark data
6 configurations
| Hardware | Framework | Quant | Decode | VRAM used | Samples | Evidence |
|---|---|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell | vLLM | IQ3_S | 9.3 tok/s | 96 GB | 1 | L0 Self-reported |
| 2x NVIDIA RTX 5090 | vLLM | IQ2_XS | — | 64 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | Strata | IQ2_XXS | 7.8 tok/s | 18 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5080 SUPER | llama.cpp | IQ2_XS | 6.4 tok/s | 24 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | vLLM | IQ2_XS | 11.1 tok/s | 32 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | llama.cpp | IQ2_XXS | 6.8 tok/s | 24 GB | 1 | L0 Self-reported |
6 of these records are not linked to a specific quantized build; the full set is on the search page.
Community discussion
No discussions yet — start the first one.
Start a discussion