Mistral Large 3 70B · Q4_K_M
L0 自报Apple M5 Max · 128GB · llama.cpp
其他
今天
11.3 tok/s
Decode
84.8 tok/s
Prefill
1.18 s
TTFT
43 GB
VRAM
配置(最近一次实测)
来源:其他 · 1 次实测 · 1 个独立来源
计数为运行配置组合(模型 + 硬件 + 框架 + 量化的聚合),同一组合的多次实测已合并;原始次数见各结果的来源行。
Apple M5 Max · 128GB · llama.cpp
其他
今天
11.3 tok/s
Decode
84.8 tok/s
Prefill
1.18 s
TTFT
43 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
Apple M4 Max · 128GB · llama.cpp
其他
今天
18 tok/s
Decode
135 tok/s
Prefill
0.74 s
TTFT
24.1 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 SUPER · 18GB · Strata
其他
今天
38 tok/s
Decode
285 tok/s
Prefill
0.35 s
TTFT
18 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5080 SUPER · 24GB · llama.cpp
其他
今天
20.6 tok/s
Decode
154.5 tok/s
Prefill
0.65 s
TTFT
24 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5090 · 32GB · vLLM
其他
今天
35.8 tok/s
Decode
268.5 tok/s
Prefill
0.37 s
TTFT
32 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 4090 · 24GB · llama.cpp
其他
今天
18.5 tok/s
Decode
138.8 tok/s
Prefill
0.72 s
TTFT
24 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
共 6 条结果