Mistral Large series 3rd gen; multilingual, long context, enhanced tool use.
0
Quantized builds
6
Benchmark records
Quantized builds
Community and official quantized releases; "linked records" counts only benchmarks tied to that exact release.
No quantized builds cataloged yet.
Benchmark data
6 configurations
| Hardware | Framework | Quant | Decode | VRAM used | Samples | Evidence |
|---|---|---|---|---|---|---|
| Apple M5 Max | llama.cpp | Q4_K_M | 11.3 tok/s | 43 GB | 1 | L0 Self-reported |
| Apple M4 Max | llama.cpp | IQ2_XS | 18 tok/s | 24.1 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | Strata | IQ2_XS | 38 tok/s | 18 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5080 SUPER | llama.cpp | Q4_K_M | 20.6 tok/s | 24 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | vLLM | Q4_K_M | 35.8 tok/s | 32 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | llama.cpp | Q4_K_M | 18.5 tok/s | 24 GB | 1 | L0 Self-reported |
6 of these records are not linked to a specific quantized build; the full set is on the search page.
Community discussion
No discussions yet — start the first one.
Start a discussion