gemma-4-12B-it · Q8_0
L0 自报NVIDIA RTX 5060 SUPER · 12GB · vLLM
其他
今天
29.3 tok/s
Decode
219.8 tok/s
Prefill
0.45 s
TTFT
12 GB
VRAM
配置(最近一次实测)
来源:其他 · 1 次实测 · 1 个独立来源
计数为运行配置组合(模型 + 硬件 + 框架 + 量化的聚合),同一组合的多次实测已合并;原始次数见各结果的来源行。
NVIDIA RTX 5060 SUPER · 12GB · vLLM
其他
今天
29.3 tok/s
Decode
219.8 tok/s
Prefill
0.45 s
TTFT
12 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5060 SUPER · 12GB · vLLM
其他
今天
42.8 tok/s
Decode
321 tok/s
Prefill
0.31 s
TTFT
9.4 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5060 SUPER · 12GB · vLLM
其他
今天
52.2 tok/s
Decode
391.5 tok/s
Prefill
0.26 s
TTFT
7.8 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
其他
今天
26.9 tok/s
Decode
201.8 tok/s
Prefill
0.5 s
TTFT
12 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
其他
今天
39.3 tok/s
Decode
294.8 tok/s
Prefill
0.34 s
TTFT
9.4 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
其他
今天
47.9 tok/s
Decode
359.3 tok/s
Prefill
0.28 s
TTFT
7.8 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 SUPER · 18GB · vLLM
其他
今天
58.6 tok/s
Decode
439.5 tok/s
Prefill
0.23 s
TTFT
13.5 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 SUPER · 18GB · vLLM
其他
今天
85.6 tok/s
Decode
642 tok/s
Prefill
0.16 s
TTFT
9.4 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 SUPER · 18GB · vLLM
其他
今天
104.3 tok/s
Decode
782.3 tok/s
Prefill
0.13 s
TTFT
7.8 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 SUPER · 18GB · llama.cpp
其他
今天
53.8 tok/s
Decode
403.5 tok/s
Prefill
0.25 s
TTFT
13.5 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 SUPER · 18GB · llama.cpp
其他
今天
78.5 tok/s
Decode
588.8 tok/s
Prefill
0.17 s
TTFT
9.4 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 SUPER · 18GB · llama.cpp
其他
今天
95.7 tok/s
Decode
717.8 tok/s
Prefill
0.14 s
TTFT
7.8 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 3090 · 24GB · vLLM
其他
今天
61.2 tok/s
Decode
459 tok/s
Prefill
0.22 s
TTFT
13.5 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 3090 · 24GB · vLLM
其他
今天
89.4 tok/s
Decode
670.5 tok/s
Prefill
0.15 s
TTFT
9.4 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 3090 · 24GB · vLLM
其他
今天
109 tok/s
Decode
817.5 tok/s
Prefill
0.12 s
TTFT
7.8 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 3090 · 24GB · llama.cpp
其他
今天
56.2 tok/s
Decode
421.5 tok/s
Prefill
0.24 s
TTFT
13.5 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 3090 · 24GB · llama.cpp
其他
今天
82 tok/s
Decode
615 tok/s
Prefill
0.16 s
TTFT
9.4 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 3090 · 24GB · llama.cpp
其他
今天
100 tok/s
Decode
750 tok/s
Prefill
0.13 s
TTFT
7.8 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 Ti · 16GB · vLLM
其他
今天
58.6 tok/s
Decode
439.5 tok/s
Prefill
0.23 s
TTFT
13.5 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
NVIDIA RTX 5070 Ti · 16GB · vLLM
其他
今天
85.6 tok/s
Decode
642 tok/s
Prefill
0.16 s
TTFT
9.4 GB
VRAM
来源:其他 · 1 次实测 · 1 个独立来源
共 42 条结果