Llama 5 70B

70BChatby Meta Start a discussion

Llama 5 series mid-size dense; 70B parameters, native multimodal (text + image + audio).

0

Quantized builds

6

Benchmark records

Quantized builds

Community and official quantized releases; "linked records" counts only benchmarks tied to that exact release.

No quantized builds cataloged yet.

Benchmark data

6 configurations

HardwareFrameworkQuantDecodeVRAM usedSamplesEvidence
Apple M5 Maxllama.cppQ4_K_M11.3 tok/s43 GB1L0 Self-reported
Apple M4 Maxllama.cppIQ2_XS18 tok/s24.1 GB1L0 Self-reported
NVIDIA RTX 5070 SUPERStrataIQ2_XS38 tok/s18 GB1L0 Self-reported
NVIDIA RTX 5080 SUPERllama.cppQ5_K_M16.8 tok/s24 GB1L0 Self-reported
NVIDIA RTX 5090vLLMQ4_K_M35.8 tok/s32 GB1L0 Self-reported
NVIDIA RTX 4090llama.cppQ4_K_M18.5 tok/s24 GB1L0 Self-reported

6 of these records are not linked to a specific quantized build; the full set is on the search page.

Community discussion

No discussions yet — start the first one.

Start a discussion