Apple M5 Ultra
DesktopApple's workstation-class unified-memory SoC, launched 2026-08-25 in the new Mac Studio: up to 36 CPU cores, 80 GPU cores, 512 GB of unified memory and 1.2 TB/s (about 1200 GB/s) of memory bandwidth — 50% higher than the previous generation. Apple does not publish a TDP, so that column is left empty.
19
Run records
7
Models covered
1
Frameworks covered
L0 Self-reported × 19
Evidence mix
Typical measured performance
Published benchmark averages per model × framework × quantization on this hardware.
- llama.cpp · IQ2_XS Decode 3.7 · Prefill 27.8 ·
- llama.cpp · IQ3_S Decode 3.1 · Prefill 23.3 ·
- llama.cpp · IQ2_XS Decode 3.6 · Prefill 27 ·
- llama.cpp · IQ3_S Decode 3 · Prefill 22.5 ·
- llama.cpp · IQ2_XS Decode 102.9 · Prefill 771.8 ·
- llama.cpp · Q5_K_M Decode 46.8 · Prefill 351 ·
- llama.cpp · Q4_K_M Decode 57.1 · Prefill 428.3 ·
- llama.cpp · Q4_K_M Decode 22 · Prefill 165 ·
- llama.cpp · IQ3_S Decode 33 · Prefill 247.5 ·
- llama.cpp · IQ2_XS Decode 39.7 · Prefill 297.8 ·
- llama.cpp · Q4_K_M Decode 22 · Prefill 165 ·
- llama.cpp · IQ3_S Decode 33 · Prefill 247.5 ·
- llama.cpp · IQ2_XS Decode 39.7 · Prefill 297.8 ·
- llama.cpp · Q4_K_M Decode 21.4 · Prefill 160.5 ·
- llama.cpp · IQ3_S Decode 32.1 · Prefill 240.8 ·
- llama.cpp · IQ2_XS Decode 38.5 · Prefill 288.8 ·
- llama.cpp · Q8_0 Decode 108.8 · Prefill 816 ·
- llama.cpp · Q5_K_M Decode 158.6 · Prefill 1189.5 ·
- llama.cpp · Q4_K_M Decode 191 · Prefill 1432.5 ·
Key specs
- Vendor
- Apple
- Architecture
- M5 Ultra
- VRAM
- 512 GB
- Memory bandwidth
- 1200 GB/s
- Release date
- 2026-08-25
Run records
19 configurations
| Model | Framework | Quantization | Decode | Gen time | Samples |
|---|---|---|---|---|---|
| GLM-5.3 | llama.cpp | IQ2_XS | 3.7 tok/s | — | 1 |
| GLM-5.3 | llama.cpp | IQ3_S | 3.1 tok/s | — | 1 |
| DeepSeek-V4.1-Flash | llama.cpp | IQ2_XS | 3.6 tok/s | — | 1 |
| DeepSeek-V4.1-Flash | llama.cpp | IQ3_S | 3 tok/s | — | 1 |
| Qwen3.8-27B | llama.cpp | IQ2_XS | 102.9 tok/s | — | 1 |
| Qwen3.8-27B | llama.cpp | Q5_K_M | 46.8 tok/s | — | 1 |
| Qwen3.8-27B | llama.cpp | Q4_K_M | 57.1 tok/s | — | 1 |
| Mistral Large 3 70B | llama.cpp | Q4_K_M | 22 tok/s | — | 1 |
| Mistral Large 3 70B | llama.cpp | IQ3_S | 33 tok/s | — | 1 |
| Mistral Large 3 70B | llama.cpp | IQ2_XS | 39.7 tok/s | — | 1 |
| Llama 5 70B | llama.cpp | Q4_K_M | 22 tok/s | — | 1 |
| Llama 5 70B | llama.cpp | IQ3_S | 33 tok/s | — | 1 |
| Llama 5 70B | llama.cpp | IQ2_XS | 39.7 tok/s | — | 1 |
| Qwen 4 72B | llama.cpp | Q4_K_M | 21.4 tok/s | — | 1 |
| Qwen 4 72B | llama.cpp | IQ3_S | 32.1 tok/s | — | 1 |
| Qwen 4 72B | llama.cpp | IQ2_XS | 38.5 tok/s | — | 1 |
| Qwen 4 8B | llama.cpp | Q8_0 | 108.8 tok/s | — | 1 |
| Qwen 4 8B | llama.cpp | Q5_K_M | 158.6 tok/s | — | 1 |
| Qwen 4 8B | llama.cpp | Q4_K_M | 191 tok/s | — | 1 |