Benchmark search

Try these searches

6 matching benchmark runs

Hardware: AllModel: AllFramework: AllQuantization: AllEvidence level: AllSource: All

Qwen3-32B · Q4_K_M

L2 Cross-framework verified

NVIDIA RTX 4090 · 24GB · llama.cpp

GitHub

2 days ago

18.4 tok/s

Decode

216.93 tok/s

Prefill

1.19 s

TTFT

19.5 GB

VRAM

Config (latest run):

Ubuntu 24.04 CUDA 12.4 Q4_K_M llama.cpp b4567 Flash Attention

Sources: GitHub, Reddit · 3 runs · 2 independent sources

Qwen3-32B · Q4_K_M

L1 Reproduced

NVIDIA RTX 4090 · 24GB · Ollama

V2EX

2 days ago

17.4 tok/s

Decode

198.2 tok/s

Prefill

1.31 s

TTFT

19.8 GB

VRAM

Config (latest run):

Ubuntu 24.04 CUDA 12.4 Q4_K_M Ollama 0.5.7

Sources: V2EX · 1 runs · 1 independent sources

Qwen3-32B · Q4_K_M

L1 Reproduced

NVIDIA RTX 3090 · 24GB · llama.cpp

GitHub

2 days ago

12.95 tok/s

Decode

143.35 tok/s

Prefill

1.6 s

TTFT

19.5 GB

VRAM

Config (latest run):

Ubuntu 22.04 CUDA 12.2 Q4_K_M llama.cpp b4567

Sources: GitHub, Bilibili · 2 runs · 2 independent sources

Llama-3.1-8B-Instruct · Q5_K_M

L1 Reproduced

NVIDIA RTX 4060 Ti 16GB · 16GB · llama.cpp

GitHub

2 days ago

33.8 tok/s

Decode

620.4 tok/s

Prefill

0.41 s

TTFT

6.2 GB

VRAM

Config (latest run):

Windows 11 24H2 CUDA 12.4 Q5_K_M llama.cpp b4602 Flash Attention

Sources: GitHub · 1 runs · 1 independent sources

Llama-3.1-8B-Instruct · Q4_K_M

L1 Reproduced

Apple M3 Max(MacBook Pro 14) · 36GB · Ollama

Reddit

2 days ago

39.15 tok/s

Decode

408.85 tok/s

Prefill

0.51 s

TTFT

5.8 GB

VRAM

Config (latest run):

macOS 15.3 Q4_K_M Ollama 0.5.7

Sources: Reddit, Hugging Face · 2 runs · 2 independent sources

DeepSeek-R1-Distill-Qwen-14B · AWQ-INT4

L2 Cross-framework verified

NVIDIA RTX 4090 · 24GB · vLLM

GitHub

2 days ago

63.6 tok/s

Decode

1,192.75 tok/s

Prefill

0.29 s

TTFT

11.1 GB

VRAM

Config (latest run):

Ubuntu 24.04 CUDA 12.4 AWQ-INT4 vLLM 0.7.3

Sources: GitHub, Hugging Face · 2 runs · 2 independent sources

6 results