News
Open-source LLM 5-way comparison: Qwen 3.x / DeepSeek V4.1 / GLM-5.3 / Llama 4.5 / Gemma 4
5-way open-source flagship comparison based on local DB RTX 4090 measurements: Qwen 3.5 27B 91 tok/s, Gemma 4 12B 89 tok/s, Qwen 3.8 Flash 180B 12 tok/s, DeepSeek V4.1 763B 3 tok/s. 27B is the sweet spot.
Source: araoai
Open-source LLM 5-way comparison: Qwen 3.x / DeepSeek V4.1 / GLM-5.3 / Llama 4.5 / Gemma 4
Five open-source flagships — Qwen / DeepSeek / GLM / Llama / Gemma. This post is based on local DB measurements + vendor announcements.
Five families
| Vendor | Model | Params | License |
|---|---|---|---|
| Alibaba Qwen | Qwen3.8-Flash-Next | 180B | OPEN_WEIGHTS |
| Alibaba Qwen | Qwen3.5-27B | 27B | UNVERIFIED |
| DeepSeek | DeepSeek-V4.1-Flash | 763B | OPEN_SOURCE |
| DeepSeek | DeepSeek-V4-Flash | 284B | OPEN_SOURCE |
| Z.ai GLM | GLM-5.3 | 753B | OPEN_WEIGHTS |
| Z.ai GLM | GLM-5.3-Flash | 320B | OPEN_SOURCE |
| Meta Llama | Llama 4.5 Maverick | 400B+ | Llama Community |
| Google Gemma | gemma-4-31B-it | 31B | UNVERIFIED |
| Google Gemma | gemma-4-12B-it | 12B | UNVERIFIED |
RTX 4090 measured decode tok/s
| Model | Params | tok/s |
|---|---|---|
| Qwen3.5-27B | 27B | 91.25 |
| gemma-4-12B-it | 12B | 89.35 |
| gemma-4-31B-it | 31B | 52.57 |
| Qwen3.8-Flash-Next | 180B | 11.60 |
| DeepSeek-V4-Flash | 284B | 5.77 |
| GLM-5.3-Flash | 320B | 5.25 |
| GLM-5.3 | 753B | 3.40 |
| DeepSeek-V4.1-Flash | 763B | 3.30 |
5-axis rating (1-5 stars)
| Dimension | Qwen 3.x | DeepSeek V4 | GLM-5.3 | Llama 4.5 | Gemma 4 |
|---|---|---|---|---|---|
| Closes-source parity | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐ |
| Local-runnable (≤30B) | ⭐⭐⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐⭐⭐ |
| License cleanness | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐⭐⭐ |
| Ecosystem support | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐ |
| Roadmap clarity | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐ | ⭐⭐⭐⭐ |
Build recommendations
| Your card | Pick |
|---|---|
| 12GB (3060/5070) | Qwen 3.5 27B (91 tok/s equivalent) |
| 24GB (4090/3090) | Gemma 4 31B (52 tok/s) |
| 32GB+ single (5090) | Future GLM-5.3 / DeepSeek V4.1 quantized |
| Multi 4×24GB | Qwen3.8-Flash-Next 11.6 tok/s |
| Multi 8×24GB | DeepSeek V4.1 Flash (3.3 tok/s, long context) |
Three judgment points
- 700B+ models nearly unusable on a single card — 3 tok/s is poor
- Llama 4.5 has no local DB data — Llama section based on vendor announcements (not measured)
- "Qwen 3.7" is colloquial — actual DB models are Qwen 3.5/3.6/3.8 series
Evidence
- 91.25 / 89.35 / 52.57 / 11.60 / 5.77 / 5.25 / 3.40 / 3.30 tok/s: local DB benchmark_record (hardware_id=1 = RTX 4090)
- 5-axis rating: combined DB data + vendor announcements + third-party reports
- "Meta may not ship next gen": Alibaba Apsara 10/2 + multiple recaps (uncertain event)
Sources
- Local DB: benchmark_record table + model table
- Llama 4.5: Meta official announcement (not measured)