Qwen3.6-27B · Q5_K_S
L0 Self-reportedNVIDIA RTX 3090 · 24GB · llama.cpp
GitHub
Today
46.57 tok/s
Decode
1,201.22 tok/s
Prefill
0.1 s
TTFT
21.53 GB
VRAM
Config (latest run):
Debian GNU/Linux 13 (trixie) CUDA 13.3 Q5_K_S llama.cpp v0.3.0-e0663be2713c
Sources: GitHub · 1 runs · 1 independent sources