2026
Rethinking Convergence in MoE Training: The Role of Routing Sparsity
ICML 2026poster
In Mixture-of-Experts (MoE) training, sparse routing, i.e., activating only the top-$K$ experts per token, is essential for balancing convergence speed and computational cost. However, existing works typically choose $K$ empirically, without theoretical guidance. To address this gap, we characterize…