← 新闻
动态

开源大模型 5 家横评:Qwen 3.x / DeepSeek V4.1 / GLM-5.3 / Llama 4.5 / Gemma 4

5 家开源旗舰横评,基于本地 DB RTX 4090 实测:Qwen 3.5 27B 91 tok/s、Gemma 4 12B 89 tok/s、Qwen 3.8 Flash 180B 12 tok/s、DeepSeek V4.1 763B 3 tok/s。27B 是单卡甜蜜点。

来源: araoai

开源大模型 5 家横评:Qwen 3.x / DeepSeek V4.1 / GLM-5.3 / Llama 4.5 / Gemma 4

开源阵营现在 5 家——Qwen / DeepSeek / GLM / Llama / Gemma。本帖基于本地 DB 实测 + 厂商公告横评。

5 家基本盘

厂商 代表型号 参数 许可证
Alibaba Qwen Qwen3.8-Flash-Next 180B OPEN_WEIGHTS
Alibaba Qwen Qwen3.5-27B 27B UNVERIFIED
DeepSeek DeepSeek-V4.1-Flash 763B OPEN_SOURCE
DeepSeek DeepSeek-V4-Flash 284B OPEN_SOURCE
Z.ai GLM GLM-5.3 753B OPEN_WEIGHTS
Z.ai GLM GLM-5.3-Flash 320B OPEN_SOURCE
Meta Llama Llama 4.5 Maverick 400B+ Llama Community
Google Gemma gemma-4-31B-it 31B UNVERIFIED
Google Gemma gemma-4-12B-it 12B UNVERIFIED

RTX 4090 实测 decode tok/s

模型 参数 tok/s
Qwen3.5-27B 27B 91.25
gemma-4-12B-it 12B 89.35
gemma-4-31B-it 31B 52.57
Qwen3.8-Flash-Next 180B 11.60
DeepSeek-V4-Flash 284B 5.77
GLM-5.3-Flash 320B 5.25
GLM-5.3 753B 3.40
DeepSeek-V4.1-Flash 763B 3.30

5 维评分(1-5 星)

维度 Qwen 3.x DeepSeek V4 GLM-5.3 Llama 4.5 Gemma 4
能力追平闭源 ⭐⭐⭐⭐ ⭐⭐⭐⭐⭐ ⭐⭐⭐⭐ ⭐⭐⭐⭐ ⭐⭐⭐
本地跑得动(≤30B) ⭐⭐⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐⭐⭐
许可证干净 ⭐⭐⭐⭐ ⭐⭐⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐⭐⭐
生态工具支持 ⭐⭐⭐⭐⭐ ⭐⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐ ⭐⭐⭐
长期路线图清晰 ⭐⭐⭐⭐ ⭐⭐⭐⭐⭐ ⭐⭐⭐ ⭐⭐ ⭐⭐⭐⭐

装机推荐

你的卡 推荐组合
12GB 显存(3060/5070) Qwen 3.5 27B(91 tok/s 等效)
24GB 显存(4090/3090) Gemma 4 31B(52 tok/s)
32GB+ 单卡(5090) 未来 GLM-5.3 / DeepSeek V4.1 量化版
多卡 4×24GB Qwen3.8-Flash-Next 11.6 tok/s
多卡 8×24GB DeepSeek V4.1 Flash(3.3 tok/s,长上下文用)

三个判断点

  1. 700B+ 模型在单卡几乎不可用——3 tok/s 体验差
  2. Llama 4.5 在我 DB 还没数据——本帖 Llama 部分基于厂商公告(未实测)
  3. Qwen "3.7" 是口语化指代——DB 里实际是 Qwen3.5/3.6/3.8 三个系列

证据分层

  • 91.25 / 89.35 / 52.57 / 11.60 / 5.77 / 5.25 / 3.40 / 3.30 tok/s:本地 DB benchmark_record(hardware_id=1 = RTX 4090)
  • 5 维评分:综合 DB 数据 + 厂商公告 + 第三方报告
  • "Meta 可能不再推下一代":Alibaba Apsara 10/2 + 多家复盘(不确定事件)

数据来源

  • 本地 DB:benchmark_record 表 + model 表
  • Llama 4.5 部分:Meta 官方公告(未实测)