← News
News

Open-source LLM 5-way comparison: Qwen 3.x / DeepSeek V4.1 / GLM-5.3 / Llama 4.5 / Gemma 4

5-way open-source flagship comparison based on local DB RTX 4090 measurements: Qwen 3.5 27B 91 tok/s, Gemma 4 12B 89 tok/s, Qwen 3.8 Flash 180B 12 tok/s, DeepSeek V4.1 763B 3 tok/s. 27B is the sweet spot.

Source: araoai

Open-source LLM 5-way comparison: Qwen 3.x / DeepSeek V4.1 / GLM-5.3 / Llama 4.5 / Gemma 4

Five open-source flagships — Qwen / DeepSeek / GLM / Llama / Gemma. This post is based on local DB measurements + vendor announcements.

Five families

Vendor Model Params License
Alibaba Qwen Qwen3.8-Flash-Next 180B OPEN_WEIGHTS
Alibaba Qwen Qwen3.5-27B 27B UNVERIFIED
DeepSeek DeepSeek-V4.1-Flash 763B OPEN_SOURCE
DeepSeek DeepSeek-V4-Flash 284B OPEN_SOURCE
Z.ai GLM GLM-5.3 753B OPEN_WEIGHTS
Z.ai GLM GLM-5.3-Flash 320B OPEN_SOURCE
Meta Llama Llama 4.5 Maverick 400B+ Llama Community
Google Gemma gemma-4-31B-it 31B UNVERIFIED
Google Gemma gemma-4-12B-it 12B UNVERIFIED

RTX 4090 measured decode tok/s

Model Params tok/s
Qwen3.5-27B 27B 91.25
gemma-4-12B-it 12B 89.35
gemma-4-31B-it 31B 52.57
Qwen3.8-Flash-Next 180B 11.60
DeepSeek-V4-Flash 284B 5.77
GLM-5.3-Flash 320B 5.25
GLM-5.3 753B 3.40
DeepSeek-V4.1-Flash 763B 3.30

5-axis rating (1-5 stars)

Dimension Qwen 3.x DeepSeek V4 GLM-5.3 Llama 4.5 Gemma 4
Closes-source parity ⭐⭐⭐⭐ ⭐⭐⭐⭐⭐ ⭐⭐⭐⭐ ⭐⭐⭐⭐ ⭐⭐⭐
Local-runnable (≤30B) ⭐⭐⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐⭐⭐
License cleanness ⭐⭐⭐⭐ ⭐⭐⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐⭐⭐
Ecosystem support ⭐⭐⭐⭐⭐ ⭐⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐
Roadmap clarity ⭐⭐⭐⭐ ⭐⭐⭐⭐⭐ ⭐⭐⭐ ⭐⭐ ⭐⭐⭐⭐

Build recommendations

Your card Pick
12GB (3060/5070) Qwen 3.5 27B (91 tok/s equivalent)
24GB (4090/3090) Gemma 4 31B (52 tok/s)
32GB+ single (5090) Future GLM-5.3 / DeepSeek V4.1 quantized
Multi 4×24GB Qwen3.8-Flash-Next 11.6 tok/s
Multi 8×24GB DeepSeek V4.1 Flash (3.3 tok/s, long context)

Three judgment points

  1. 700B+ models nearly unusable on a single card — 3 tok/s is poor
  2. Llama 4.5 has no local DB data — Llama section based on vendor announcements (not measured)
  3. "Qwen 3.7" is colloquial — actual DB models are Qwen 3.5/3.6/3.8 series

Evidence

  • 91.25 / 89.35 / 52.57 / 11.60 / 5.77 / 5.25 / 3.40 / 3.30 tok/s: local DB benchmark_record (hardware_id=1 = RTX 4090)
  • 5-axis rating: combined DB data + vendor announcements + third-party reports
  • "Meta may not ship next gen": Alibaba Apsara 10/2 + multiple recaps (uncertain event)

Sources

  • Local DB: benchmark_record table + model table
  • Llama 4.5: Meta official announcement (not measured)