Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.
ModelQwen3.8-9B
Qwen3.8-9B · Q4_K_M
L0 Self-reported
Intel Arc B580 · 12GB · llama.cpp
Other
Today
— tok/s
Decode
— tok/s
Prefill
— s
TTFT
6 GB
VRAM
Config (latest run)
Linux CUDA 12.6 Q4_K_M llama.cpp b4200 Flash Attention