Qwen3-1.7B · IQ2_M
L2 跨框架验证NVIDIA RTX 4090 · 24GB · llama.cpp
其他
今天
323.97 tok/s
Decode
1,188.09 tok/s
Prefill
2.05 s
TTFT
3.96 GB
VRAM
配置(最近一次实测)
来源:其他 · 1 次实测 · 1 个独立来源
计数为运行配置组合(模型 + 硬件 + 框架 + 量化的聚合),同一组合的多次实测已合并;原始次数见各结果的来源行。
NVIDIA RTX 4090 · 24GB · llama.cpp
其他
今天
323.97 tok/s
Decode
1,188.09 tok/s
Prefill
2.05 s
TTFT
3.96 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 4090 · 24GB · llama.cpp
其他
今天
306.65 tok/s
Decode
1,436.24 tok/s
Prefill
1.31 s
TTFT
4.06 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 4090 · 24GB · Strata
其他
今天
287.87 tok/s
Decode
1,066.63 tok/s
Prefill
0.32 s
TTFT
3.6 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 4090 · 24GB · Strata
其他
今天
285.75 tok/s
Decode
1,116.76 tok/s
Prefill
0.6 s
TTFT
4.06 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 3090 · 24GB · llama.cpp
其他
今天
227.45 tok/s
Decode
862.55 tok/s
Prefill
0.46 s
TTFT
4.67 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 3090 · 24GB · llama.cpp
其他
今天
175.42 tok/s
Decode
578.57 tok/s
Prefill
0.86 s
TTFT
5.31 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 3090 · 24GB · Strata
其他
今天
215.56 tok/s
Decode
902.66 tok/s
Prefill
0.24 s
TTFT
5.67 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
Apple M3 Max · 36GB · vLLM
其他
今天
309.15 tok/s
Decode
1,109.54 tok/s
Prefill
2.17 s
TTFT
3.81 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
Apple M3 Max · 36GB · vLLM
其他
今天
226.47 tok/s
Decode
1,116.37 tok/s
Prefill
0.18 s
TTFT
4.67 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
Apple M3 Max · 36GB · Ollama
其他
今天
96.85 tok/s
Decode
510.37 tok/s
Prefill
0.44 s
TTFT
8.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
Apple M3 Max · 36GB · Ollama
其他
今天
188.36 tok/s
Decode
704.89 tok/s
Prefill
1.31 s
TTFT
4.75 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
Apple M3 Max · 36GB · Strata
其他
今天
239.13 tok/s
Decode
849.25 tok/s
Prefill
1.36 s
TTFT
5.31 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
AMD Radeon RX 7900 XTX · 24GB · llama.cpp
其他
今天
234.47 tok/s
Decode
1,079.52 tok/s
Prefill
2.39 s
TTFT
4.32 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
AMD Radeon RX 7900 XTX · 24GB · llama.cpp
其他
今天
178.72 tok/s
Decode
842.7 tok/s
Prefill
0.98 s
TTFT
5.03 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
AMD Radeon RX 7900 XTX · 24GB · Ollama
其他
今天
196.24 tok/s
Decode
705.31 tok/s
Prefill
0.21 s
TTFT
4.53 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
AMD Radeon RX 7900 XTX · 24GB · Ollama
其他
今天
111.35 tok/s
Decode
362.65 tok/s
Prefill
0.89 s
TTFT
6.21 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
AMD Radeon RX 7900 XTX · 24GB · Strata
其他
今天
249 tok/s
Decode
1,211.72 tok/s
Prefill
2.07 s
TTFT
4.88 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5090 · 32GB · llama.cpp
其他
今天
353.77 tok/s
Decode
1,906.09 tok/s
Prefill
1.04 s
TTFT
5.42 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5090 · 32GB · llama.cpp
其他
今天
345.13 tok/s
Decode
1,779.25 tok/s
Prefill
0.42 s
TTFT
5.67 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5090 · 32GB · Ollama
其他
今天
199.27 tok/s
Decode
1,057.11 tok/s
Prefill
2.17 s
TTFT
8.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
共 41 条结果