Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.
Modelgemma-4-E2B-it
gemma-4-E2B-it · Q4_K_M
L0 Self-reported
NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
Other
Today
— tok/s
Decode
— tok/s
Prefill
— s
TTFT
— GB
VRAM
Config (latest run)
Linux CUDA 12.6 Q4_K_M llama.cpp b4200 Flash Attention