Qwen3.5-27B · Q4_K_M
L1 ReproducedNVIDIA RTX 5090 · 32GB · llama.cpp
GitHub
Yesterday
30.09 tok/s
Decode
1,515.83 tok/s
Prefill
— s
TTFT
— GB
VRAM
Config (latest run)
Windows 11 Q4_K_M llama.cpp b10741
Sources: GitHub · 1 runs · 1 independent sources