Benchmark search

Try these searches

1 matching benchmark runs

ModelGLM-5.3-FlashHardware: AllFramework: AllQuantization: AllEvidence level: AllSource: All

GLM-5.3-Flash · UD-IQ4_XS

L0 Self-reported

2x NVIDIA RTX 5090 · 64GB · llama.cpp

GitHub

Today

21.96 tok/s

Decode

421.07 tok/s

Prefill

0.6 s

TTFT

60.3 GB

VRAM

Config (latest run):

Ubuntu 24.04.4 LTS CUDA 13.3 UD-IQ4_XS llama.cpp moecachev1.6-rc0

Sources: GitHub · 1 runs · 1 independent sources

1 results