Qwen3.6-35B-A3B · IQ2_XS
L2 Cross-framework verifiedApple M3 Max · 36GB · llama.cpp
Other
Today
89.41 tok/s
Decode
406.89 tok/s
Prefill
0.17 s
TTFT
34.2 GB
VRAM
Config (latest run)
Sources: Other · 1 runs · 1 independent sources
Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.
Apple M3 Max · 36GB · llama.cpp
Other
Today
89.41 tok/s
Decode
406.89 tok/s
Prefill
0.17 s
TTFT
34.2 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
Apple M3 Max · 36GB · vLLM
Other
Today
102.56 tok/s
Decode
411.25 tok/s
Prefill
1.72 s
TTFT
32.97 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
Apple M4 Max · 128GB · llama.cpp
Other
Today
121.63 tok/s
Decode
540.05 tok/s
Prefill
2.37 s
TTFT
47.67 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
Apple M4 Max · 128GB · llama.cpp
Other
Today
111.01 tok/s
Decode
387.61 tok/s
Prefill
2.26 s
TTFT
59.43 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
Apple M4 Max · 128GB · Ollama
Other
Today
102.3 tok/s
Decode
395.71 tok/s
Prefill
2.39 s
TTFT
55.02 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
5 results