Qwen 4 72B · Q4_K_M
L0 自报Apple M5 Max · 128GB · llama.cpp
其他
今天
11 tok/s
Decode
82.5 tok/s
Prefill
1.21 s
TTFT
44.2 GB
VRAM
配置(最近一次实测)
来源:其他 · 1 次实测 · 1 个独立来源
计数为运行配置组合(模型 + 硬件 + 框架 + 量化的聚合),同一组合的多次实测已合并;原始次数见各结果的来源行。
Apple M5 Max · 128GB · llama.cpp
其他
今天
11 tok/s
Decode
82.5 tok/s
Prefill
1.21 s
TTFT
44.2 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
Apple M4 Max · 128GB · llama.cpp
其他
今天
17.5 tok/s
Decode
131.3 tok/s
Prefill
0.76 s
TTFT
24.8 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 SUPER · 18GB · Strata
其他
今天
36.9 tok/s
Decode
276.8 tok/s
Prefill
0.36 s
TTFT
18 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5080 SUPER · 24GB · llama.cpp
其他
今天
20 tok/s
Decode
150 tok/s
Prefill
0.67 s
TTFT
24 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5090 · 32GB · vLLM
其他
今天
34.9 tok/s
Decode
261.8 tok/s
Prefill
0.38 s
TTFT
32 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 4090 · 24GB · llama.cpp
其他
今天
18 tok/s
Decode
135 tok/s
Prefill
0.74 s
TTFT
24 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
共 6 条结果