Qwen3.5-35B-A3B · Q4_K_M
L0 Self-reportedApple M5 Max · 128GB · llama.cpp
Other
Today
21.9 tok/s
Decode
164.3 tok/s
Prefill
0.61 s
TTFT
22.4 GB
VRAM
Config (latest run)
Sources: Other · 1 runs · 1 independent sources
Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.
Apple M5 Max · 128GB · llama.cpp
Other
Today
21.9 tok/s
Decode
164.3 tok/s
Prefill
0.61 s
TTFT
22.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5080 SUPER · 24GB · llama.cpp
Other
Today
32.7 tok/s
Decode
245.3 tok/s
Prefill
0.41 s
TTFT
24 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5090 · 32GB · vLLM
Other
Today
69.6 tok/s
Decode
522 tok/s
Prefill
0.19 s
TTFT
22.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 4090 · 24GB · llama.cpp
Other
Today
35.9 tok/s
Decode
269.3 tok/s
Prefill
0.37 s
TTFT
22.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
4 results