Apple M3 Ultra
DesktopApple's workstation-class unified-memory SoC (Mac Studio, shipping 2025-03-12): up to 32 CPU cores and 80 GPU cores, unified memory starting at 96 GB and configurable to 512 GB, with 819 GB/s of memory bandwidth. Apple's own framing is that it can run LLMs with more than 600B parameters entirely in memory. Apple does not publish a TDP, so that column is left empty.
15
Run records
5
Models covered
1
Frameworks covered
L0 Self-reported × 15
Evidence mix
Typical measured performance
Published benchmark averages per model × framework × quantization on this hardware.
- llama.cpp · IQ2_XS Decode 70.2 · Prefill 526.5 ·
- llama.cpp · Q5_K_M Decode 31.9 · Prefill 239.3 ·
- llama.cpp · Q4_K_M Decode 39 · Prefill 292.5 ·
- llama.cpp · Q4_K_M Decode 15 · Prefill 112.5 ·
- llama.cpp · IQ3_S Decode 22.5 · Prefill 168.8 ·
- llama.cpp · IQ2_XS Decode 27.1 · Prefill 203.3 ·
- llama.cpp · Q4_K_M Decode 15 · Prefill 112.5 ·
- llama.cpp · IQ3_S Decode 22.5 · Prefill 168.8 ·
- llama.cpp · IQ2_XS Decode 27.1 · Prefill 203.3 ·
- llama.cpp · Q4_K_M Decode 14.6 · Prefill 109.5 ·
- llama.cpp · IQ3_S Decode 21.9 · Prefill 164.3 ·
- llama.cpp · IQ2_XS Decode 26.3 · Prefill 197.3 ·
- llama.cpp · Q8_0 Decode 74.3 · Prefill 557.3 ·
- llama.cpp · Q5_K_M Decode 108.3 · Prefill 812.3 ·
- llama.cpp · Q4_K_M Decode 130.4 · Prefill 978 ·
Key specs
- Vendor
- Apple
- Architecture
- M3 Ultra
- VRAM
- 512 GB
- Memory bandwidth
- 819 GB/s
- Release date
- 2025-03-12
Run records
15 configurations
| Model | Framework | Quantization | Decode | Gen time | Samples |
|---|---|---|---|---|---|
| Qwen3.8-27B | llama.cpp | IQ2_XS | 70.2 tok/s | — | 1 |
| Qwen3.8-27B | llama.cpp | Q5_K_M | 31.9 tok/s | — | 1 |
| Qwen3.8-27B | llama.cpp | Q4_K_M | 39 tok/s | — | 1 |
| Mistral Large 3 70B | llama.cpp | Q4_K_M | 15 tok/s | — | 1 |
| Mistral Large 3 70B | llama.cpp | IQ3_S | 22.5 tok/s | — | 1 |
| Mistral Large 3 70B | llama.cpp | IQ2_XS | 27.1 tok/s | — | 1 |
| Llama 5 70B | llama.cpp | Q4_K_M | 15 tok/s | — | 1 |
| Llama 5 70B | llama.cpp | IQ3_S | 22.5 tok/s | — | 1 |
| Llama 5 70B | llama.cpp | IQ2_XS | 27.1 tok/s | — | 1 |
| Qwen 4 72B | llama.cpp | Q4_K_M | 14.6 tok/s | — | 1 |
| Qwen 4 72B | llama.cpp | IQ3_S | 21.9 tok/s | — | 1 |
| Qwen 4 72B | llama.cpp | IQ2_XS | 26.3 tok/s | — | 1 |
| Qwen 4 8B | llama.cpp | Q8_0 | 74.3 tok/s | — | 1 |
| Qwen 4 8B | llama.cpp | Q5_K_M | 108.3 tok/s | — | 1 |
| Qwen 4 8B | llama.cpp | Q4_K_M | 130.4 tok/s | — | 1 |