Llama 5 70B · Q4_K_M
L0 Self-reportedApple M5 Max · 128GB · llama.cpp
Other
Today
11.3 tok/s
Decode
84.8 tok/s
Prefill
1.18 s
TTFT
43 GB
VRAM
Config (latest run)
Sources: Other · 1 runs · 1 independent sources
Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.
Apple M5 Max · 128GB · llama.cpp
Other
Today
11.3 tok/s
Decode
84.8 tok/s
Prefill
1.18 s
TTFT
43 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
Apple M4 Max · 128GB · llama.cpp
Other
Today
18 tok/s
Decode
135 tok/s
Prefill
0.74 s
TTFT
24.1 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5070 SUPER · 18GB · Strata
Other
Today
38 tok/s
Decode
285 tok/s
Prefill
0.35 s
TTFT
18 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5080 SUPER · 24GB · llama.cpp
Other
Today
16.8 tok/s
Decode
126 tok/s
Prefill
0.79 s
TTFT
24 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 5090 · 32GB · vLLM
Other
Today
35.8 tok/s
Decode
268.5 tok/s
Prefill
0.37 s
TTFT
32 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
NVIDIA RTX 4090 · 24GB · llama.cpp
Other
Today
18.5 tok/s
Decode
138.8 tok/s
Prefill
0.72 s
TTFT
24 GB
VRAM
Sources: Other · 1 runs · 1 independent sources
6 results