← News
News

Why I am bullish on the Qwen team: 3 things — model power + ecosystem + team power

Qwen camp three pillars: model power (Qwen 3.7 235B flagship position) + ecosystem (Strata prioritizes Qwen 4) + team power (ID Balancing paper 89.6% improvement). Qwen 4 still "in training".

Source: araoai

Why I am bullish on the Qwen team: 3 things — model power + ecosystem + team power

In the past 4 days I wrote 3 Qwen-related posts — strung together they are my judgment for "still holding" the Qwen camp.

1️⃣ Model power (April open vs closed gap)

Thorsten Meyer's April piece "Single Digits": open vs closed on mainstream benchmarks narrowed from 20-30% to single digits (<10%).

Qwen camp standing:

Model Vendor Type Status
Qwen 3.7 235B Alibaba Apache 2.0 H1 flagship; Qwen 4 in training
DeepSeek V4.1 Flash DeepSeek MIT 1.6T / 49B active
GLM-5.3 Z.ai MIT Anthropic called "most cyber-capable"
Llama 4.5 Maverick Meta Llama Community Meta may not ship next gen
Gemma 4 31B Google Apache 2.0 Clean license

Key judgment: Qwen 3.7 235B is the current "open-source flagship" position — single-machine runnable (multi-card sharding), Apache 2.0 weights, capability on par with Claude Opus 5.5 / GPT-6.1 Sol.

2️⃣ Ecosystem (frontier inference tools prioritize Qwen 4)

Strata author Niko Veit 2026-10-06 tweet:

"Strata will support Qwen4 fully... I can bring same performance as you now have on 5090 or 5070 with just 6GB VRAM."

Niko is targeting Qwen 4 as the first full-adaptation target — not empty words. Strata GitHub current model list has Qwen 3.x / 3.5 / 3.6 / 3.7 / 3.8 fully supported; Qwen 4 = v0.6.

Key judgment: frontier inference frameworks prioritize Qwen = Qwen is now the "de-facto standard candidate" on the LLM inference side. Meta not shipping next gen = ecosystem loses a major player, Qwen fills in.

3️⃣ Team power (ID Balancing paper)

arXiv 2609.39137v1: Alibaba Qwen team recasts MoE expert load balancing as a PID control problem.

Method Control view Worst MaxVio (Top-3-of-768)
Aux loss baseline Traditional gradient penalty ~150
DeepSeek loss-free Pure I (integral) 211.19
Kimi K3 Quantile Pure P (proportional) 31.11
ID Balancing (Qwen team) I + D 15.33

Across scales 18.9B → 69.9B, 89.6% improvement vs aux loss baseline.

Key judgment: Qwen team doesn't just ship models, they also ship training methodology papers. This is what Meta / DeepSeek / Mistral are doing — Qwen is keeping up.

Three things connected

Dimension Evidence How I use it
Model power Qwen 3.7 235B Apache 2.0 Local primary (27B runnable + 235B multi-card fallback)
Ecosystem Strata prioritizes Qwen 4 No plan to change card, wait for Strata v0.6
Team power ID Balancing 89.6% improvement Watch Qwen team output for 6.0 roadmap

Three judgment points

  1. Qwen 4 still "in training" (Alibaba Apsara 10/2). All my Qwen 4 references are forward-looking claims
  2. Meta not shipping next gen if true, open-source camp loses a major player
  3. ID Balancing paper is engineering view, not new architecture

Evidence

  • April open vs closed gap: thorstenmeyerai.com H1 recap
  • Qwen 3.7 flagship: Alibaba official + mainstream benchmark recap
  • Strata prioritizes Qwen 4: Niko Veit @coldniko 2026-10-06 tweet
  • ID Balancing 89.6% improvement: arXiv 2609.39137v1
  • "Meta may not ship next gen": Alibaba Apsara 10/2 + multiple recaps

Sources

  • April open vs closed gap: https://thorstenmeyerai.com/insights/single-digits-the-april-that-closed-the-open-weight-gap
  • Strata prioritizes Qwen 4: https://x.com/coldniko/status/2107406312682258511
  • ID Balancing paper: https://arxiv.org/html/2609.39137v1