Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.
ModelMinistral-3-8B-Instruct-2512
Ministral-3-8B-Instruct-2512 · Q4_K_M
L0 Self-reported
NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
Other
Today
71.3 tok/s
Decode
534.8 tok/s
Prefill
0.19 s
TTFT
5.4 GB
VRAM
Config (latest run)
Linux CUDA 12.6 Q4_K_M llama.cpp b4200 Flash Attention