Intel Arc A770 16GB
GPU16
Run records
1
Models covered
4
Frameworks covered
L2 Cross-framework verified × 16
Evidence mix
Typical measured performance
Published benchmark averages per model × framework × quantization on this hardware.
- llama.cpp · IQ4_XS Decode 84 · Prefill 777 ·
- llama.cpp · Q4_K_M Decode 77 · Prefill 681 ·
- llama.cpp · Q5_K_M Decode 68 · Prefill 739 ·
- llama.cpp · Q8_0 Decode 47 · Prefill 559 ·
- vLLM · IQ4_XS Decode 69 · Prefill 706 ·
- vLLM · Q4_K_M Decode 62 · Prefill 655 ·
- vLLM · Q5_K_M Decode 54 · Prefill 646 ·
- vLLM · Q8_0 Decode 45 · Prefill 434 ·
- Ollama · IQ4_XS Decode 66 · Prefill 791 ·
- Ollama · Q4_K_M Decode 62 · Prefill 744 ·
- Ollama · Q5_K_M Decode 68 · Prefill 680 ·
- Ollama · Q8_0 Decode 50 · Prefill 494 ·
- Strata · IQ4_XS Decode 115 · Prefill 1228 ·
- Strata · Q4_K_M Decode 105 · Prefill 1087 ·
- Strata · Q5_K_M Decode 100 · Prefill 950 ·
- Strata · Q8_0 Decode 63 · Prefill 752 ·
Key specs
- Vendor
- Intel
- VRAM
- 16 GB
Run records
16 configurations
| Model | Framework | Quantization | Decode | Gen time | Samples |
|---|---|---|---|---|---|
| Qwen3.5-27B | llama.cpp | IQ4_XS | 84 tok/s | — | 1 |
| Qwen3.5-27B | llama.cpp | Q4_K_M | 77 tok/s | — | 1 |
| Qwen3.5-27B | llama.cpp | Q5_K_M | 68 tok/s | — | 1 |
| Qwen3.5-27B | llama.cpp | Q8_0 | 47 tok/s | — | 1 |
| Qwen3.5-27B | vLLM | IQ4_XS | 69 tok/s | — | 1 |
| Qwen3.5-27B | vLLM | Q4_K_M | 62 tok/s | — | 1 |
| Qwen3.5-27B | vLLM | Q5_K_M | 54 tok/s | — | 1 |
| Qwen3.5-27B | vLLM | Q8_0 | 45 tok/s | — | 1 |
| Qwen3.5-27B | Ollama | IQ4_XS | 66 tok/s | — | 1 |
| Qwen3.5-27B | Ollama | Q4_K_M | 62 tok/s | — | 1 |
| Qwen3.5-27B | Ollama | Q5_K_M | 68 tok/s | — | 1 |
| Qwen3.5-27B | Ollama | Q8_0 | 50 tok/s | — | 1 |
| Qwen3.5-27B | Strata | IQ4_XS | 115 tok/s | — | 1 |
| Qwen3.5-27B | Strata | Q4_K_M | 105 tok/s | — | 1 |
| Qwen3.5-27B | Strata | Q5_K_M | 100 tok/s | — | 1 |
| Qwen3.5-27B | Strata | Q8_0 | 63 tok/s | — | 1 |