[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f1zcuk50kmwswg":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":21,"bodyZh":28,"bodyEn":29,"sourceName":30,"sourceUrl":31,"updatedAt":32},"id-balancing-pid-moe-qwen-2026-10-09","RUNTIME","2026-10-09T12:00:00.000Z","阿里 Qwen 团队论文：ID Balancing——MoE 训练专家负载不均的 PID 解法","Alibaba Qwen team paper: ID Balancing — PID-based load control for sparse MoE training","arXiv 2609.39137v1 论文：把 MoE 专家负载平衡重新表述为 PID 控制问题。Top-3-of-768 MaxVio 从 211 → 15.33，跨规模 18.9B→69.9B 89.6% 改善 vs aux loss。","arXiv 2609.39137v1: MoE expert load balancing recast as PID control. Top-3-of-768 MaxVio 211→15.33, 89.6% improvement vs aux loss across 18.9B→69.9B.",[12,15,18],{"id":13,"name":14},5,"Qwen3.8-27B",{"id":16,"name":17},6,"Qwen3.6-27B",{"id":19,"name":20},50,"Qwen3.5-27B",[22,25],{"id":23,"name":24},1,"NVIDIA RTX 4090",{"id":26,"name":27},55,"NVIDIA H100 80GB HBM3","# 阿里 Qwen 团队论文：ID Balancing——MoE 训练专家负载不均的 PID 解法\n\narXiv 2609.39137v1 — \"ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control\"\n\n## 问题（MoE 训练老毛病）\n\nMoE 模型路由越稀疏，每个 token 选中的 expert 越少 → 专家分配越敏感 → 一些 expert 过载、一些无人问津。效率掉、稳定掉、参数利用率掉。\n\n之前三大流派：\n- **Aux loss**：加梯度惩罚（影响主 loss、收敛变慢）\n- **DeepSeek loss-free**：用 bias 调负载（简单但负载大时跟不上）\n- **Kimi K3 Quantile Balancing**：用分位数做平衡（平滑但反应慢）\n\n## 论文核心：把平衡问题重新表述为 PID 控制\n\nPID = 比例（P）+ 积分（I）+ 微分（D）—— 自动控制经典框架\n\n| 方法 | 控制视角 | 行为 |\n|---|---|---|\n| DeepSeek loss-free | 纯 I（积分）| 固定步长累加负载误差 |\n| Kimi K3 Quantile | 纯 P（比例）| 用分位数估计即时偏差 |\n| **ID Balancing** | **I + D** | 累加误差 + 偏差变坏时加速响应 |\n\n**关键洞察**：P 处理\"现在有多偏\"，I 处理\"过去累积偏多少\"，D 处理\"偏差在变好还是变坏\"——只在变坏时介入（derivative gate）。\n\n## 公式（每层 O(E)）\n\n```\ne_i = (n_avg - n_i) \u002F n_avg   \u002F\u002F 归一化负载误差\nb_i ← b_i + K_i · e_i + K_d · g_i · (e_i - e_i_prev)\ng_i = 1 当 e_i 与 (e_i - e_i_prev) 同号\ng_i = 0 否则\nb ← b - mean(b)              \u002F\u002F 重新中心化\n```\n\n工程细节：\n- 只用 token-count 反馈，**O(E) 每层**，不引入新梯度\n- 不用 aux loss——主 loss 不被打扰\n- K_i = K_d = 6e-3（pretrain）；续训时降到 6e-6 修漂移\n\n## 实测\n\n| 方法 | Worst Backbone MaxVio（Top-3-of-768，18.9B \u002F 120B tokens）|\n|---|---|\n| Aux loss baseline | ~150 |\n| DeepSeek loss-free | 211.19 |\n| Kimi K3 Quantile | 31.11 |\n| **ID Balancing** | **15.33** |\n\n跨规模 18.9B → 69.9B，**89.6% 改善** vs aux loss baseline。\n\n## 关键扩展\n\n- 跨规模稳定：18.9B → 69.9B MaxVio 几乎不变\n- 2.3× 学习率仍稳：aux loss 和 loss-free 在 2.3× LR 下显著退化\n- 下游质量不打折：24.8B 模型 9 个 benchmark 平均 55.81；69.9B 模型 58.66\n\n## 证据分层\n\n- 211 → 15.33 MaxVio：论文正文 Table 1\n- 89.6% vs aux loss：69.9B 规模 Top-10-of-768 验证\n- PID 视角统一三种方法：论文 §1 + §2 公式推导\n- 9 个下游 benchmark：MMLU, MMLU-Pro, SuperGPQA, MATH, GSM8K, BBH, MMMLU, EvalPlus, MultiPL-E\n\n## 数据来源\n\n- 论文：https:\u002F\u002Farxiv.org\u002Fhtml\u002F2609.39137v1\n- 作者：Peng Jin, Zihan Qiu, Zekun Wang, Bo Zheng, Yang Xu, Tian Xie, Xiao Li, Huaqing Zhang, Haoran Lian, Rui Men, Dayiheng Liu（Alibaba Qwen Team）","# Alibaba Qwen team paper: ID Balancing — PID-based load control for sparse MoE training\n\narXiv 2609.39137v1 — \"ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control\"\n\n## Problem (old MoE training pain)\n\nAs MoE models route fewer experts per token, expert assignments become more sensitive: some experts overloaded, some unused. Efficiency drops, stability drops, parameter utilization drops.\n\nThree prior schools:\n- **Aux loss**: gradient penalty (pollutes main loss, slows convergence)\n- **DeepSeek loss-free**: bias-based load adjustment (simple but can't keep up under heavy load)\n- **Kimi K3 Quantile Balancing**: quantile-based balance (smooth but slow to react)\n\n## Core: recast balancing as PID control\n\nPID = Proportional (P) + Integral (I) + Derivative (D) — classical control theory\n\n| Method | Control view | Behavior |\n|---|---|---|\n| DeepSeek loss-free | Pure I (integral) | Fixed-step cumulative load error |\n| Kimi K3 Quantile | Pure P (proportional) | Quantile-based instant bias estimate |\n| **ID Balancing** | **I + D** | Cumulative error + accelerated response when bias worsens |\n\n**Key insight**: P handles \"how biased now\", I handles \"cumulative bias over time\", D handles \"is bias improving or worsening\" — only intervenes when worsening (derivative gate).\n\n## Formula (O(E) per layer)\n\n```\ne_i = (n_avg - n_i) \u002F n_avg\nb_i ← b_i + K_i · e_i + K_d · g_i · (e_i - e_i_prev)\ng_i = 1 if e_i and (e_i - e_i_prev) have same sign\ng_i = 0 otherwise\nb ← b - mean(b)  \u002F\u002F re-center\n```\n\nEngineering details:\n- Only token-count feedback, **O(E) per layer**, no new gradients\n- No aux loss — main loss undisturbed\n- K_i = K_d = 6e-3 (pretrain); 6e-6 in continued pretraining for drift correction\n\n## Results\n\n| Method | Worst Backbone MaxVio (Top-3-of-768, 18.9B \u002F 120B tokens) |\n|---|---|\n| Aux loss baseline | ~150 |\n| DeepSeek loss-free | 211.19 |\n| Kimi K3 Quantile | 31.11 |\n| **ID Balancing** | **15.33** |\n\nAcross scales 18.9B → 69.9B, **89.6% improvement** vs aux loss baseline.\n\n## Key extensions\n\n- Scale-stable: MaxVio nearly unchanged from 18.9B to 69.9B\n- 2.3× learning rate still stable: aux loss and loss-free degrade significantly at 2.3× LR\n- Downstream quality preserved: 24.8B model averages 55.81 across 9 benchmarks; 69.9B model 58.66\n\n## Evidence\n\n- 211 → 15.33 MaxVio: paper Table 1\n- 89.6% vs aux loss: 69.9B scale Top-10-of-768 validation\n- PID view unifies three methods: paper §1 + §2 formula derivation\n- 9 downstream benchmarks: MMLU, MMLU-Pro, SuperGPQA, MATH, GSM8K, BBH, MMMLU, EvalPlus, MultiPL-E\n\n## Sources\n\n- Paper: https:\u002F\u002Farxiv.org\u002Fhtml\u002F2609.39137v1\n- Authors: Peng Jin, Zihan Qiu, Zekun Wang, Bo Zheng, Yang Xu, Tian Xie, Xiao Li, Huaqing Zhang, Haoran Lian, Rui Men, Dayiheng Liu (Alibaba Qwen Team)","arXiv","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2609.39137v1","2026-10-10T05:41:10.782Z"]