Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.
ModelQwen3.5-0.8B
Qwen3.5-0.8B · Q8_0
L0 Self-reported
NVIDIA RTX 5060 SUPER · 12GB · llama.cpp
Other
Today
388.3 tok/s
Decode
2,912.3 tok/s
Prefill
0.03 s
TTFT
1.4 GB
VRAM
Config (latest run)
Linux CUDA 12.6 Q8_0 llama.cpp b4200 Flash Attention