GLM-5.3-Flash · UD-IQ4_XS
L0 Self-reported2x NVIDIA RTX 5090 · 64GB · llama.cpp
GitHub
Today
21.96 tok/s
Decode
421.07 tok/s
Prefill
0.6 s
TTFT
60.3 GB
VRAM
Config (latest run):
Ubuntu 24.04.4 LTS CUDA 13.3 UD-IQ4_XS llama.cpp moecachev1.6-rc0
Sources: GitHub · 1 runs · 1 independent sources