Qwen 4 72B

72BChatby Qwen (Alibaba) Start a discussion

Qwen 4 series mid-size dense; 72B parameters, strong reasoning + code + long context.

0

Quantized builds

6

Benchmark records

Quantized builds

Community and official quantized releases; "linked records" counts only benchmarks tied to that exact release.

No quantized builds cataloged yet.

Benchmark data

6 configurations

HardwareFrameworkQuantDecodeVRAM usedSamplesEvidence
Apple M5 Maxllama.cppQ4_K_M11 tok/s44.2 GB1L0 Self-reported
Apple M4 Maxllama.cppIQ2_XS17.5 tok/s24.8 GB1L0 Self-reported
NVIDIA RTX 5070 SUPERStrataIQ2_XS36.9 tok/s18 GB1L0 Self-reported
NVIDIA RTX 5080 SUPERllama.cppQ4_K_M20 tok/s24 GB1L0 Self-reported
NVIDIA RTX 5090vLLMQ4_K_M34.9 tok/s32 GB1L0 Self-reported
NVIDIA RTX 4090llama.cppQ4_K_M18 tok/s24 GB1L0 Self-reported

6 of these records are not linked to a specific quantized build; the full set is on the search page.

Community discussion

No discussions yet — start the first one.

Start a discussion