Llama-3.1-8B-Instruct · Q4_K_M
This page aggregates 2 real-world runs of Llama-3.1-8B-Instruct (Q4_K_M) on Apple M3 Max(MacBook Pro 14) with Ollama, contributed by 2 independent source platforms; metrics are averages of published measurements.
39.15 tok/s
Decode
Decode speed
408.85 tok/s
Prefill
Prefill speed
0.51 s
TTFT
Time to first token
5.8 GB
VRAM
VRAM usage
2
Measured runs
2
Independent sources
Reddit, Hugging Face
Source platforms
2026-08-18
Last verified
Performance
- Ollama · Q4_K_M (current)Decode 39.15 · Prefill 408.85 ·
Core figures
- Decode (avg)
- 39.15 tok/s
- Prefill (avg)
- 408.85 tok/s
- TTFT (avg)
- 0.51 s
- VRAM (avg)
- 5.8 GB
- MTP acceptance rate
- —
- TTFB
- — GB
- Power draw
- 47.3 W
Configuration
Member-level fields are taken from the most recent run
- Model
- Llama-3.1-8B-Instruct
- Quantization
- Q4_K_M
- Framework
- Ollama
- Version
- 0.5.7
- Context length
- 8192 tokens
Hardware
Nominal and measured figures are shown side by side; whether it runs is the reader's call
- GPU
- Apple M3 Max(MacBook Pro 14)
- Nominal VRAM
- 36 GB
- Measured VRAM (avg)
- 5.8 GB
- OS
- macOS 15.3
Sources & evidence
2 measured records in total, each traceable to its original source
0.5.7 · macOS 15.3 · 8192 ctx
38.9 tok/s
Decode
402.1 tok/s
Prefill
0.52 s
TTFT
5.8 GB
VRAM
—
MTP
— W
Power draw
M3 Max 统一内存下 8B Q4 约 39 tok/s,功耗约 45W。
0.5.9 · macOS 15.3 · 8192 ctx
39.4 tok/s
Decode
415.6 tok/s
Prefill
0.49 s
TTFT
5.8 GB
VRAM
—
MTP
47.3 W
Power draw
Ollama 0.5.9 复现 39.4 tok/s。