Why I am bullish on the Qwen team: 3 things — model power + ecosystem + team power
Qwen camp three pillars: model power (Qwen 3.7 235B flagship position) + ecosystem (Strata prioritizes Qwen 4) + team power (ID Balancing paper 89.6% improvement). Qwen 4 still "in training".
Source: araoai
Why I am bullish on the Qwen team: 3 things — model power + ecosystem + team power
In the past 4 days I wrote 3 Qwen-related posts — strung together they are my judgment for "still holding" the Qwen camp.
1️⃣ Model power (April open vs closed gap)
Thorsten Meyer's April piece "Single Digits": open vs closed on mainstream benchmarks narrowed from 20-30% to single digits (<10%).
Qwen camp standing:
| Model | Vendor | Type | Status |
|---|---|---|---|
| Qwen 3.7 235B | Alibaba | Apache 2.0 | H1 flagship; Qwen 4 in training |
| DeepSeek V4.1 Flash | DeepSeek | MIT | 1.6T / 49B active |
| GLM-5.3 | Z.ai | MIT | Anthropic called "most cyber-capable" |
| Llama 4.5 Maverick | Meta | Llama Community | Meta may not ship next gen |
| Gemma 4 31B | Apache 2.0 | Clean license |
Key judgment: Qwen 3.7 235B is the current "open-source flagship" position — single-machine runnable (multi-card sharding), Apache 2.0 weights, capability on par with Claude Opus 5.5 / GPT-6.1 Sol.
2️⃣ Ecosystem (frontier inference tools prioritize Qwen 4)
Strata author Niko Veit 2026-10-06 tweet:
"Strata will support Qwen4 fully... I can bring same performance as you now have on 5090 or 5070 with just 6GB VRAM."
Niko is targeting Qwen 4 as the first full-adaptation target — not empty words. Strata GitHub current model list has Qwen 3.x / 3.5 / 3.6 / 3.7 / 3.8 fully supported; Qwen 4 = v0.6.
Key judgment: frontier inference frameworks prioritize Qwen = Qwen is now the "de-facto standard candidate" on the LLM inference side. Meta not shipping next gen = ecosystem loses a major player, Qwen fills in.
3️⃣ Team power (ID Balancing paper)
arXiv 2609.39137v1: Alibaba Qwen team recasts MoE expert load balancing as a PID control problem.
| Method | Control view | Worst MaxVio (Top-3-of-768) |
|---|---|---|
| Aux loss baseline | Traditional gradient penalty | ~150 |
| DeepSeek loss-free | Pure I (integral) | 211.19 |
| Kimi K3 Quantile | Pure P (proportional) | 31.11 |
| ID Balancing (Qwen team) | I + D | 15.33 |
Across scales 18.9B → 69.9B, 89.6% improvement vs aux loss baseline.
Key judgment: Qwen team doesn't just ship models, they also ship training methodology papers. This is what Meta / DeepSeek / Mistral are doing — Qwen is keeping up.
Three things connected
| Dimension | Evidence | How I use it |
|---|---|---|
| Model power | Qwen 3.7 235B Apache 2.0 | Local primary (27B runnable + 235B multi-card fallback) |
| Ecosystem | Strata prioritizes Qwen 4 | No plan to change card, wait for Strata v0.6 |
| Team power | ID Balancing 89.6% improvement | Watch Qwen team output for 6.0 roadmap |
Three judgment points
- Qwen 4 still "in training" (Alibaba Apsara 10/2). All my Qwen 4 references are forward-looking claims
- Meta not shipping next gen if true, open-source camp loses a major player
- ID Balancing paper is engineering view, not new architecture
Evidence
- April open vs closed gap: thorstenmeyerai.com H1 recap
- Qwen 3.7 flagship: Alibaba official + mainstream benchmark recap
- Strata prioritizes Qwen 4: Niko Veit @coldniko 2026-10-06 tweet
- ID Balancing 89.6% improvement: arXiv 2609.39137v1
- "Meta may not ship next gen": Alibaba Apsara 10/2 + multiple recaps
Sources
- April open vs closed gap: https://thorstenmeyerai.com/insights/single-digits-the-april-that-closed-the-open-weight-gap
- Strata prioritizes Qwen 4: https://x.com/coldniko/status/2107406312682258511
- ID Balancing paper: https://arxiv.org/html/2609.39137v1