Qwen3-32B · Q4_K_M
L2 Cross-framework verifiedNVIDIA RTX 4090 · 24GB · llama.cpp
GitHub
2 days ago
18.4 tok/s
Decode
216.93 tok/s
Prefill
1.19 s
TTFT
19.5 GB
VRAM
Config (latest run):
Sources: GitHub, Reddit · 3 runs · 2 independent sources
NVIDIA RTX 4090 · 24GB · llama.cpp
GitHub
2 days ago
18.4 tok/s
Decode
216.93 tok/s
Prefill
1.19 s
TTFT
19.5 GB
VRAM
Config (latest run):
Sources: GitHub, Reddit · 3 runs · 2 independent sources
NVIDIA RTX 4090 · 24GB · Ollama
V2EX
2 days ago
17.4 tok/s
Decode
198.2 tok/s
Prefill
1.31 s
TTFT
19.8 GB
VRAM
Config (latest run):
Sources: V2EX · 1 runs · 1 independent sources
NVIDIA RTX 3090 · 24GB · llama.cpp
GitHub
2 days ago
12.95 tok/s
Decode
143.35 tok/s
Prefill
1.6 s
TTFT
19.5 GB
VRAM
Config (latest run):
Sources: GitHub, Bilibili · 2 runs · 2 independent sources
NVIDIA RTX 4060 Ti 16GB · 16GB · llama.cpp
GitHub
2 days ago
33.8 tok/s
Decode
620.4 tok/s
Prefill
0.41 s
TTFT
6.2 GB
VRAM
Config (latest run):
Sources: GitHub · 1 runs · 1 independent sources
Apple M3 Max(MacBook Pro 14) · 36GB · Ollama
2 days ago
39.15 tok/s
Decode
408.85 tok/s
Prefill
0.51 s
TTFT
5.8 GB
VRAM
Config (latest run):
Sources: Reddit, Hugging Face · 2 runs · 2 independent sources
NVIDIA RTX 4090 · 24GB · vLLM
GitHub
2 days ago
63.6 tok/s
Decode
1,192.75 tok/s
Prefill
0.29 s
TTFT
11.1 GB
VRAM
Config (latest run):
Sources: GitHub, Hugging Face · 2 runs · 2 independent sources
6 results