Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.
ModelQwen3-4B
Qwen3-4B · Q4_K_M
L0 Self-reported
NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
Other
Today
145.6 tok/s
Decode
1,092 tok/s
Prefill
0.09 s
TTFT
2.9 GB
VRAM
Config (latest run)
Linux CUDA 12.6 Q4_K_M llama.cpp b4200 Flash Attention