Qwen 4 8B · IQ4_XS
L0 Self-reportedNVIDIA RTX 5070 SUPER · 18GB · Strata
Other
Today
208.4 tok/s
Decode
1,563 tok/s
Prefill
0.06 s
TTFT
4.8 GB
VRAM
Config (latest run)
Sources: Other · 1 runs · 1 independent sources
Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.
NVIDIA RTX 5070 SUPER · 18GB · Strata
Other
Today
208.4 tok/s
Decode
1,563 tok/s
Prefill
0.06 s
TTFT
4.8 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5080 SUPER · 24GB · llama.cpp
Other
Today
178.3 tok/s
Decode
1,337.3 tok/s
Prefill
0.07 s
TTFT
5.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5090 · 32GB · vLLM
Other
Today
258.2 tok/s
Decode
1,936.5 tok/s
Prefill
0.05 s
TTFT
6.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 4090 · 24GB · llama.cpp
Other
Today
160.5 tok/s
Decode
1,203.8 tok/s
Prefill
0.08 s
TTFT
5.4 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
4 results