Alibaba Qwen team paper: ID Balancing — PID-based load control for sparse MoE training
arXiv 2609.39137v1: MoE expert load balancing recast as PID control. Top-3-of-768 MaxVio 211→15.33, 89.6% improvement vs aux loss across 18.9B→69.9B.
Source: arXiv
Alibaba Qwen team paper: ID Balancing — PID-based load control for sparse MoE training
arXiv 2609.39137v1 — "ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control"
Problem (old MoE training pain)
As MoE models route fewer experts per token, expert assignments become more sensitive: some experts overloaded, some unused. Efficiency drops, stability drops, parameter utilization drops.
Three prior schools:
- Aux loss: gradient penalty (pollutes main loss, slows convergence)
- DeepSeek loss-free: bias-based load adjustment (simple but can't keep up under heavy load)
- Kimi K3 Quantile Balancing: quantile-based balance (smooth but slow to react)
Core: recast balancing as PID control
PID = Proportional (P) + Integral (I) + Derivative (D) — classical control theory
| Method | Control view | Behavior |
|---|---|---|
| DeepSeek loss-free | Pure I (integral) | Fixed-step cumulative load error |
| Kimi K3 Quantile | Pure P (proportional) | Quantile-based instant bias estimate |
| ID Balancing | I + D | Cumulative error + accelerated response when bias worsens |
Key insight: P handles "how biased now", I handles "cumulative bias over time", D handles "is bias improving or worsening" — only intervenes when worsening (derivative gate).
Formula (O(E) per layer)
e_i = (n_avg - n_i) / n_avg
b_i ← b_i + K_i · e_i + K_d · g_i · (e_i - e_i_prev)
g_i = 1 if e_i and (e_i - e_i_prev) have same sign
g_i = 0 otherwise
b ← b - mean(b) // re-center
Engineering details:
- Only token-count feedback, O(E) per layer, no new gradients
- No aux loss — main loss undisturbed
- K_i = K_d = 6e-3 (pretrain); 6e-6 in continued pretraining for drift correction
Results
| Method | Worst Backbone MaxVio (Top-3-of-768, 18.9B / 120B tokens) |
|---|---|
| Aux loss baseline | ~150 |
| DeepSeek loss-free | 211.19 |
| Kimi K3 Quantile | 31.11 |
| ID Balancing | 15.33 |
Across scales 18.9B → 69.9B, 89.6% improvement vs aux loss baseline.
Key extensions
- Scale-stable: MaxVio nearly unchanged from 18.9B to 69.9B
- 2.3× learning rate still stable: aux loss and loss-free degrade significantly at 2.3× LR
- Downstream quality preserved: 24.8B model averages 55.81 across 9 benchmarks; 69.9B model 58.66
Evidence
- 211 → 15.33 MaxVio: paper Table 1
- 89.6% vs aux loss: 69.9B scale Top-10-of-768 validation
- PID view unifies three methods: paper §1 + §2 formula derivation
- 9 downstream benchmarks: MMLU, MMLU-Pro, SuperGPQA, MATH, GSM8K, BBH, MMMLU, EvalPlus, MultiPL-E
Sources
- Paper: https://arxiv.org/html/2609.39137v1
- Authors: Peng Jin, Zihan Qiu, Zekun Wang, Bo Zheng, Yang Xu, Tian Xie, Xiao Li, Huaqing Zhang, Haoran Lian, Rui Men, Dayiheng Liu (Alibaba Qwen Team)