Qwen3.6-35B-A3B · IQ2_XS
L2 跨框架验证Apple M3 Max · 36GB · llama.cpp
其他
今天
89.41 tok/s
Decode
406.89 tok/s
Prefill
0.17 s
TTFT
34.2 GB
VRAM
配置(最近一次实测)
macOS 15.0 CUDA N/A IQ2_XS llama.cpp v0.25 Flash Attention
来源:其他 · 1 次实测 · 1 个独立来源
计数为运行配置组合(模型 + 硬件 + 框架 + 量化的聚合),同一组合的多次实测已合并;原始次数见各结果的来源行。
Apple M3 Max · 36GB · llama.cpp
其他
今天
89.41 tok/s
Decode
406.89 tok/s
Prefill
0.17 s
TTFT
34.2 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
Apple M3 Max · 36GB · vLLM
其他
今天
102.56 tok/s
Decode
411.25 tok/s
Prefill
1.72 s
TTFT
32.97 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
Apple M4 Max · 128GB · llama.cpp
其他
今天
121.63 tok/s
Decode
540.05 tok/s
Prefill
2.37 s
TTFT
47.67 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
Apple M4 Max · 128GB · llama.cpp
其他
今天
111.01 tok/s
Decode
387.61 tok/s
Prefill
2.26 s
TTFT
59.43 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
Apple M4 Max · 128GB · Ollama
其他
今天
102.3 tok/s
Decode
395.71 tok/s
Prefill
2.39 s
TTFT
55.02 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
共 5 条结果