Qwen 4 72B · Q4_K_M
L0 Self-reportedApple M5 Max · 128GB · llama.cpp
Other
Today
11 tok/s
Decode
82.5 tok/s
Prefill
1.21 s
TTFT
44.2 GB
VRAM
Config (latest run)
Sources: Other · 1 runs · 1 independent sources
Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.
Apple M5 Max · 128GB · llama.cpp
Other
Today
11 tok/s
Decode
82.5 tok/s
Prefill
1.21 s
TTFT
44.2 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
Apple M4 Max · 128GB · llama.cpp
Other
Today
17.5 tok/s
Decode
131.3 tok/s
Prefill
0.76 s
TTFT
24.8 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5070 SUPER · 18GB · Strata
Other
Today
36.9 tok/s
Decode
276.8 tok/s
Prefill
0.36 s
TTFT
18 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5080 SUPER · 24GB · llama.cpp
Other
Today
20 tok/s
Decode
150 tok/s
Prefill
0.67 s
TTFT
24 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5090 · 32GB · vLLM
Other
Today
34.9 tok/s
Decode
261.8 tok/s
Prefill
0.38 s
TTFT
32 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 4090 · 24GB · llama.cpp
Other
Today
18 tok/s
Decode
135 tok/s
Prefill
0.74 s
TTFT
24 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
6 results