[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f1bcf973olm5hz":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":24,"bodyZh":37,"bodyEn":38,"sourceName":39,"sourceUrl":40,"updatedAt":41},"why-bullish-on-qwen-2026-10-09","NEWS","2026-10-09T18:00:00.000Z","为什么我看好 Qwen 团队：3 件事——模型力 + 生态力 + 团队力","Why I am bullish on the Qwen team: 3 things — model power + ecosystem + team power","Qwen 阵营三个支柱：模型力（Qwen 3.7 235B 旗舰位置）+ 生态力（Strata 优先支持 Qwen 4）+ 团队力（ID Balancing 论文 89.6% 改善）。Qwen 4 仍\"训练中\"。","Qwen camp three pillars: model power (Qwen 3.7 235B flagship position) + ecosystem (Strata prioritizes Qwen 4) + team power (ID Balancing paper 89.6% improvement). Qwen 4 still \"in training\".",[12,15,18,21],{"id":13,"name":14},5,"Qwen3.8-27B",{"id":16,"name":17},6,"Qwen3.6-27B",{"id":19,"name":20},11,"Qwen3.8-Flash-Next",{"id":22,"name":23},50,"Qwen3.5-27B",[25,28,31,34],{"id":26,"name":27},1,"NVIDIA RTX 4090",{"id":29,"name":30},13,"NVIDIA RTX 5090",{"id":32,"name":33},15,"NVIDIA RTX 3060 12GB",{"id":35,"name":36},18,"NVIDIA RTX 5070","# 为什么我看好 Qwen 团队：3 件事——模型力 + 生态力 + 团队力\n\n过去 4 天写了 3 篇 Qwen 相关——串起来就是\"现在仍持有\"Qwen 阵营的判断。\n\n## 1️⃣ 模型力（4 月开源闭源差距）\n\n4 月 Thorsten Meyer 那篇《Single Digits》复盘：开源 vs 闭源在主流 benchmark 差距从 20-30% 缩到「个位数百分点」（\u003C 10%）。\n\nQwen 阵营占位：\n\n| 模型 | 厂商 | 类型 | 当前状态 |\n|---|---|---|---|\n| Qwen 3.7 235B | Alibaba | Apache 2.0 | H1 旗舰；Qwen 4 训练中 |\n| DeepSeek V4.1 Flash | DeepSeek | MIT | 1.6T \u002F 49B active |\n| GLM-5.3 | Z.ai | MIT | Anthropic 称\"最 cyber-capable\" |\n| Llama 4.5 Maverick | Meta | Llama Community | Meta 可能**不再推**下一代 |\n| Gemma 4 31B | Google | Apache 2.0 | 干净许可证 |\n\n**关键判断**：Qwen 3.7 235B 是当前\"开源旗舰\"位置——单机能跑（多卡分块）、权重 Apache 2.0、能力追平 Claude Opus 5.5 \u002F GPT-6.1 Sol。\n\n## 2️⃣ 生态力（前沿推理工具优先支持 Qwen 4）\n\nStrata 作者 Niko Veit 2026-10-06 推文：\n> \"Strata will support Qwen4 fully... I can bring same performance as you now have on 5090 or 5070 with just 6GB VRAM.\"\n\nNiko 把 Qwen 4 当**第一个全适配目标**——这不是空话。Strata GitHub 当前 model 列表里 Qwen 3.x \u002F 3.5 \u002F 3.6 \u002F 3.7 \u002F 3.8 全支持，Qwen 4 等 v0.6。\n\n**关键判断**：前沿推理框架**优先适配 Qwen** = Qwen 已是 LLM 推理侧的\"事实标准候选\"。Meta 不推下一代 = 生态少一个大型玩家，Qwen 顶上。\n\n## 3️⃣ 团队力（ID Balancing 论文）\n\narXiv 2609.39137v1 论文：阿里 Qwen 团队把 MoE 训练专家负载不均重新表述为 **PID 控制问题**。\n\n| 方法 | 控制视角 | Worst MaxVio（Top-3-of-768）|\n|---|---|---|\n| Aux loss baseline | 传统梯度惩罚 | ~150 |\n| DeepSeek loss-free | 纯 I（积分）| 211.19 |\n| Kimi K3 Quantile | 纯 P（比例）| 31.11 |\n| **ID Balancing（Qwen 团队）** | **I + D** | **15.33** |\n\n跨规模 18.9B → 69.9B，**89.6% 改善** vs aux loss baseline。\n\n**关键判断**：Qwen 团队不只发模型，**还发训练方法论文**。这是 Meta \u002F DeepSeek \u002F Mistral 都在做的事——Qwen 跟上来了。\n\n## 3 件事串联\n\n| 维度 | 证据 | 我怎么用 |\n|---|---|---|\n| 模型力 | Qwen 3.7 235B Apache 2.0 | 本地主力（27B 跑得动 + 235B 多卡分块兜底）|\n| 生态力 | Strata 优先适配 Qwen 4 | 不打算换卡等 Strata v0.6 |\n| 团队力 | ID Balancing 89.6% 改善 | 未来 6.0 路线图看 Qwen 团队输出 |\n\n## 三个判断点\n\n1. **Qwen 4 仍\"训练中\"**（Alibaba Apsara 10\u002F2 大会）。我所有 Qwen 4 引用都是 forward-looking 主张\n2. **Meta 不推下一代**这事如果成真，开源阵营少一个大型玩家\n3. **ID Balancing 论文是工程视角**，不是新架构\n\n## 证据分层\n\n- 4 月开源闭源差距：thorstenmeyerai.com 复盘 H1 数据\n- Qwen 3.7 旗舰位置：Alibaba 官方 + 主流 benchmark 复盘\n- Strata 优先支持 Qwen 4：Niko Veit @coldniko 2026-10-06 推文\n- ID Balancing 89.6% 改善：arXiv 2609.39137v1 论文\n- \"Meta 可能不再推下一代\"：Alibaba Apsara 10\u002F2 + 多家复盘\n\n## 数据来源\n\n- 4 月开源闭源差距：https:\u002F\u002Fthorstenmeyerai.com\u002Finsights\u002Fsingle-digits-the-april-that-closed-the-open-weight-gap\n- Strata 优先支持 Qwen 4：https:\u002F\u002Fx.com\u002Fcoldniko\u002Fstatus\u002F2107406312682258511\n- ID Balancing 论文：https:\u002F\u002Farxiv.org\u002Fhtml\u002F2609.39137v1","# Why I am bullish on the Qwen team: 3 things — model power + ecosystem + team power\n\nIn the past 4 days I wrote 3 Qwen-related posts — strung together they are my judgment for \"still holding\" the Qwen camp.\n\n## 1️⃣ Model power (April open vs closed gap)\n\nThorsten Meyer's April piece \"Single Digits\": open vs closed on mainstream benchmarks narrowed from 20-30% to single digits (\u003C10%).\n\nQwen camp standing:\n\n| Model | Vendor | Type | Status |\n|---|---|---|---|\n| Qwen 3.7 235B | Alibaba | Apache 2.0 | H1 flagship; Qwen 4 in training |\n| DeepSeek V4.1 Flash | DeepSeek | MIT | 1.6T \u002F 49B active |\n| GLM-5.3 | Z.ai | MIT | Anthropic called \"most cyber-capable\" |\n| Llama 4.5 Maverick | Meta | Llama Community | Meta may **not ship** next gen |\n| Gemma 4 31B | Google | Apache 2.0 | Clean license |\n\n**Key judgment**: Qwen 3.7 235B is the current \"open-source flagship\" position — single-machine runnable (multi-card sharding), Apache 2.0 weights, capability on par with Claude Opus 5.5 \u002F GPT-6.1 Sol.\n\n## 2️⃣ Ecosystem (frontier inference tools prioritize Qwen 4)\n\nStrata author Niko Veit 2026-10-06 tweet:\n> \"Strata will support Qwen4 fully... I can bring same performance as you now have on 5090 or 5070 with just 6GB VRAM.\"\n\nNiko is targeting Qwen 4 as the **first full-adaptation target** — not empty words. Strata GitHub current model list has Qwen 3.x \u002F 3.5 \u002F 3.6 \u002F 3.7 \u002F 3.8 fully supported; Qwen 4 = v0.6.\n\n**Key judgment**: frontier inference frameworks **prioritize Qwen** = Qwen is now the \"de-facto standard candidate\" on the LLM inference side. Meta not shipping next gen = ecosystem loses a major player, Qwen fills in.\n\n## 3️⃣ Team power (ID Balancing paper)\n\narXiv 2609.39137v1: Alibaba Qwen team recasts MoE expert load balancing as a **PID control problem**.\n\n| Method | Control view | Worst MaxVio (Top-3-of-768) |\n|---|---|---|\n| Aux loss baseline | Traditional gradient penalty | ~150 |\n| DeepSeek loss-free | Pure I (integral) | 211.19 |\n| Kimi K3 Quantile | Pure P (proportional) | 31.11 |\n| **ID Balancing (Qwen team)** | **I + D** | **15.33** |\n\nAcross scales 18.9B → 69.9B, **89.6% improvement** vs aux loss baseline.\n\n**Key judgment**: Qwen team doesn't just ship models, they also **ship training methodology papers**. This is what Meta \u002F DeepSeek \u002F Mistral are doing — Qwen is keeping up.\n\n## Three things connected\n\n| Dimension | Evidence | How I use it |\n|---|---|---|\n| Model power | Qwen 3.7 235B Apache 2.0 | Local primary (27B runnable + 235B multi-card fallback) |\n| Ecosystem | Strata prioritizes Qwen 4 | No plan to change card, wait for Strata v0.6 |\n| Team power | ID Balancing 89.6% improvement | Watch Qwen team output for 6.0 roadmap |\n\n## Three judgment points\n\n1. **Qwen 4 still \"in training\"** (Alibaba Apsara 10\u002F2). All my Qwen 4 references are forward-looking claims\n2. **Meta not shipping next gen** if true, open-source camp loses a major player\n3. **ID Balancing paper is engineering view**, not new architecture\n\n## Evidence\n\n- April open vs closed gap: thorstenmeyerai.com H1 recap\n- Qwen 3.7 flagship: Alibaba official + mainstream benchmark recap\n- Strata prioritizes Qwen 4: Niko Veit @coldniko 2026-10-06 tweet\n- ID Balancing 89.6% improvement: arXiv 2609.39137v1\n- \"Meta may not ship next gen\": Alibaba Apsara 10\u002F2 + multiple recaps\n\n## Sources\n\n- April open vs closed gap: https:\u002F\u002Fthorstenmeyerai.com\u002Finsights\u002Fsingle-digits-the-april-that-closed-the-open-weight-gap\n- Strata prioritizes Qwen 4: https:\u002F\u002Fx.com\u002Fcoldniko\u002Fstatus\u002F2107406312682258511\n- ID Balancing paper: https:\u002F\u002Farxiv.org\u002Fhtml\u002F2609.39137v1","araoai","https:\u002F\u002Faraoai.com\u002Fnews\u002Fwhy-bullish-on-qwen-2026-10-09","2026-10-10T05:41:10.828Z"]