Qwen3-8B · Q8_0
L0 Self-reportedNVIDIA RTX 5060 SUPER · 12GB · llama.cpp
Other
Today
40.6 tok/s
Decode
304.5 tok/s
Prefill
0.33 s
TTFT
9.1 GB
VRAM
Config (latest run)
Sources: Other · 1 runs · 1 independent sources
Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.
NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
Other
Today
40.6 tok/s
Decode
304.5 tok/s
Prefill
0.33 s
TTFT
9.1 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
Other
Today
59.2 tok/s
Decode
444 tok/s
Prefill
0.23 s
TTFT
6.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
Other
Today
71.3 tok/s
Decode
534.8 tok/s
Prefill
0.19 s
TTFT
5.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5070 SUPER · 18GB · llama.cpp
Other
Today
81.3 tok/s
Decode
609.8 tok/s
Prefill
0.16 s
TTFT
9.1 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5070 SUPER · 18GB · llama.cpp
Other
Today
118.5 tok/s
Decode
888.8 tok/s
Prefill
0.11 s
TTFT
6.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5070 SUPER · 18GB · llama.cpp
Other
Today
142.6 tok/s
Decode
1,069.5 tok/s
Prefill
0.09 s
TTFT
5.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 3090 · 24GB · llama.cpp
Other
Today
84.9 tok/s
Decode
636.8 tok/s
Prefill
0.16 s
TTFT
9.1 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 3090 · 24GB · llama.cpp
Other
Today
123.7 tok/s
Decode
927.8 tok/s
Prefill
0.11 s
TTFT
6.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 3090 · 24GB · llama.cpp
Other
Today
149 tok/s
Decode
1,117.5 tok/s
Prefill
0.09 s
TTFT
5.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5070 · 12GB · llama.cpp
Other
Today
60.9 tok/s
Decode
456.8 tok/s
Prefill
0.22 s
TTFT
9.1 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5070 · 12GB · llama.cpp
Other
Today
88.8 tok/s
Decode
666 tok/s
Prefill
0.15 s
TTFT
6.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5070 · 12GB · llama.cpp
Other
Today
107 tok/s
Decode
802.5 tok/s
Prefill
0.12 s
TTFT
5.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5090 · 32GB · llama.cpp
Other
Today
162.5 tok/s
Decode
1,218.8 tok/s
Prefill
0.08 s
TTFT
9.1 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5090 · 32GB · llama.cpp
Other
Today
236.9 tok/s
Decode
1,776.8 tok/s
Prefill
0.06 s
TTFT
6.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5090 · 32GB · llama.cpp
Other
Today
285.3 tok/s
Decode
2,139.8 tok/s
Prefill
0.05 s
TTFT
5.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 4090 · 24GB · llama.cpp
Other
Today
91.4 tok/s
Decode
685.5 tok/s
Prefill
0.15 s
TTFT
9.1 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 4090 · 24GB · llama.cpp
Other
Today
133.3 tok/s
Decode
999.8 tok/s
Prefill
0.1 s
TTFT
6.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 4090 · 24GB · llama.cpp
Other
Today
160.5 tok/s
Decode
1,203.8 tok/s
Prefill
0.08 s
TTFT
5.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
18 results