Qwen 4 720B

720BChatby Qwen (Alibaba) Start a discussion

Qwen 4 series flagship MoE; 720B total / ~70B active. 1M context, multimodal extensions.

0

Quantized builds

7

Benchmark records

Quantized builds

Community and official quantized releases; "linked records" counts only benchmarks tied to that exact release.

No quantized builds cataloged yet.

Benchmark data

7 configurations

HardwareFrameworkQuantDecodeVRAM usedSamplesEvidence
NVIDIA RTX PRO 6000 BlackwellvLLMIQ3_S5.2 tok/s96 GB1L0 Self-reported
4x NVIDIA RTX 3090vLLMIQ2_XXS—96 GB1L0 Self-reported
2x NVIDIA RTX 5090vLLMIQ2_XS—64 GB1L0 Self-reported
NVIDIA RTX 5070 SUPERStrataIQ2_XXS4.4 tok/s18 GB1L0 Self-reported
NVIDIA RTX 5080 SUPERllama.cppIQ2_XS3.6 tok/s24 GB1L0 Self-reported
NVIDIA RTX 5090vLLMIQ2_XS6.3 tok/s32 GB1L0 Self-reported
NVIDIA RTX 4090llama.cppIQ2_XS3.2 tok/s24 GB1L0 Self-reported

7 of these records are not linked to a specific quantized build; the full set is on the search page.

Community discussion

No discussions yet — start the first one.

Start a discussion