← News
Runtime

Alibaba Qwen team paper: ID Balancing — PID-based load control for sparse MoE training

arXiv 2609.39137v1: MoE expert load balancing recast as PID control. Top-3-of-768 MaxVio 211→15.33, 89.6% improvement vs aux loss across 18.9B→69.9B.

Source: arXiv

Alibaba Qwen team paper: ID Balancing — PID-based load control for sparse MoE training

arXiv 2609.39137v1 — "ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control"

Problem (old MoE training pain)

As MoE models route fewer experts per token, expert assignments become more sensitive: some experts overloaded, some unused. Efficiency drops, stability drops, parameter utilization drops.

Three prior schools:

  • Aux loss: gradient penalty (pollutes main loss, slows convergence)
  • DeepSeek loss-free: bias-based load adjustment (simple but can't keep up under heavy load)
  • Kimi K3 Quantile Balancing: quantile-based balance (smooth but slow to react)

Core: recast balancing as PID control

PID = Proportional (P) + Integral (I) + Derivative (D) — classical control theory

Method Control view Behavior
DeepSeek loss-free Pure I (integral) Fixed-step cumulative load error
Kimi K3 Quantile Pure P (proportional) Quantile-based instant bias estimate
ID Balancing I + D Cumulative error + accelerated response when bias worsens

Key insight: P handles "how biased now", I handles "cumulative bias over time", D handles "is bias improving or worsening" — only intervenes when worsening (derivative gate).

Formula (O(E) per layer)

e_i = (n_avg - n_i) / n_avg
b_i ← b_i + K_i · e_i + K_d · g_i · (e_i - e_i_prev)
g_i = 1 if e_i and (e_i - e_i_prev) have same sign
g_i = 0 otherwise
b ← b - mean(b)  // re-center

Engineering details:

  • Only token-count feedback, O(E) per layer, no new gradients
  • No aux loss — main loss undisturbed
  • K_i = K_d = 6e-3 (pretrain); 6e-6 in continued pretraining for drift correction

Results

Method Worst Backbone MaxVio (Top-3-of-768, 18.9B / 120B tokens)
Aux loss baseline ~150
DeepSeek loss-free 211.19
Kimi K3 Quantile 31.11
ID Balancing 15.33

Across scales 18.9B → 69.9B, 89.6% improvement vs aux loss baseline.

Key extensions

  • Scale-stable: MaxVio nearly unchanged from 18.9B to 69.9B
  • 2.3× learning rate still stable: aux loss and loss-free degrade significantly at 2.3× LR
  • Downstream quality preserved: 24.8B model averages 55.81 across 9 benchmarks; 69.9B model 58.66

Evidence

  • 211 → 15.33 MaxVio: paper Table 1
  • 89.6% vs aux loss: 69.9B scale Top-10-of-768 validation
  • PID view unifies three methods: paper §1 + §2 formula derivation
  • 9 downstream benchmarks: MMLU, MMLU-Pro, SuperGPQA, MATH, GSM8K, BBH, MMMLU, EvalPlus, MultiPL-E

Sources

  • Paper: https://arxiv.org/html/2609.39137v1
  • Authors: Peng Jin, Zihan Qiu, Zekun Wang, Bo Zheng, Yang Xu, Tian Xie, Xiao Li, Huaqing Zhang, Haoran Lian, Rui Men, Dayiheng Liu (Alibaba Qwen Team)