Qwen 4 series flagship MoE; 720B total / ~70B active. 1M context, multimodal extensions.
0
Quantized builds
7
Benchmark records
Quantized builds
Community and official quantized releases; "linked records" counts only benchmarks tied to that exact release.
No quantized builds cataloged yet.
Benchmark data
7 configurations
| Hardware | Framework | Quant | Decode | VRAM used | Samples | Evidence |
|---|---|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell | vLLM | IQ3_S | 5.2 tok/s | 96 GB | 1 | L0 Self-reported |
| 4x NVIDIA RTX 3090 | vLLM | IQ2_XXS | — | 96 GB | 1 | L0 Self-reported |
| 2x NVIDIA RTX 5090 | vLLM | IQ2_XS | — | 64 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5070 SUPER | Strata | IQ2_XXS | 4.4 tok/s | 18 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5080 SUPER | llama.cpp | IQ2_XS | 3.6 tok/s | 24 GB | 1 | L0 Self-reported |
| NVIDIA RTX 5090 | vLLM | IQ2_XS | 6.3 tok/s | 32 GB | 1 | L0 Self-reported |
| NVIDIA RTX 4090 | llama.cpp | IQ2_XS | 3.2 tok/s | 24 GB | 1 | L0 Self-reported |
7 of these records are not linked to a specific quantized build; the full set is on the search page.
Community discussion
No discussions yet — start the first one.
Start a discussion