Qwen2.5-Omni-3B · Q8_0
L0 自报NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
其他
今天
109.2 tok/s
Decode
819 tok/s
Prefill
0.12 s
TTFT
3.7 GB
VRAM
配置(最近一次实测)
来源:其他 · 1 次实测 · 1 个独立来源
计数为运行配置组合(模型 + 硬件 + 框架 + 量化的聚合),同一组合的多次实测已合并;原始次数见各结果的来源行。
NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
其他
今天
109.2 tok/s
Decode
819 tok/s
Prefill
0.12 s
TTFT
3.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
其他
今天
158.8 tok/s
Decode
1,191 tok/s
Prefill
0.08 s
TTFT
2.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
其他
今天
194.1 tok/s
Decode
1,455.8 tok/s
Prefill
0.07 s
TTFT
2.3 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 SUPER · 18GB · llama.cpp
其他
今天
218.4 tok/s
Decode
1,638 tok/s
Prefill
0.06 s
TTFT
3.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 SUPER · 18GB · llama.cpp
其他
今天
317.7 tok/s
Decode
2,382.8 tok/s
Prefill
0.04 s
TTFT
2.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 SUPER · 18GB · llama.cpp
其他
今天
388.3 tok/s
Decode
2,912.3 tok/s
Prefill
0.03 s
TTFT
2.3 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 3090 · 24GB · llama.cpp
其他
今天
228.2 tok/s
Decode
1,711.5 tok/s
Prefill
0.06 s
TTFT
3.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 3090 · 24GB · llama.cpp
其他
今天
331.9 tok/s
Decode
2,489.3 tok/s
Prefill
0.04 s
TTFT
2.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 3090 · 24GB · llama.cpp
其他
今天
405.6 tok/s
Decode
3,042 tok/s
Prefill
0.03 s
TTFT
2.3 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 · 12GB · llama.cpp
其他
今天
163.8 tok/s
Decode
1,228.5 tok/s
Prefill
0.08 s
TTFT
3.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 · 12GB · llama.cpp
其他
今天
238.3 tok/s
Decode
1,787.3 tok/s
Prefill
0.06 s
TTFT
2.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 · 12GB · llama.cpp
其他
今天
291.2 tok/s
Decode
2,184 tok/s
Prefill
0.05 s
TTFT
2.3 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5090 · 32GB · llama.cpp
其他
今天
436.8 tok/s
Decode
3,276 tok/s
Prefill
0.03 s
TTFT
3.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5090 · 32GB · llama.cpp
其他
今天
635.3 tok/s
Decode
4,764.8 tok/s
Prefill
0.02 s
TTFT
2.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5090 · 32GB · llama.cpp
其他
今天
776.5 tok/s
Decode
5,823.8 tok/s
Prefill
0.02 s
TTFT
2.3 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 4090 · 24GB · llama.cpp
其他
今天
245.7 tok/s
Decode
1,842.8 tok/s
Prefill
0.05 s
TTFT
3.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 4090 · 24GB · llama.cpp
其他
今天
357.4 tok/s
Decode
2,680.5 tok/s
Prefill
0.04 s
TTFT
2.7 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 4090 · 24GB · llama.cpp
其他
今天
436.8 tok/s
Decode
3,276 tok/s
Prefill
0.03 s
TTFT
2.3 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
共 18 条结果