← Search

Weihao Zhu

1 accepted papers

2026

Rethinking Convergence in MoE Training: The Role of Routing Sparsity

ICML 2026poster

In Mixture-of-Experts (MoE) training, sparse routing, i.e., activating only the top-$K$ experts per token, is essential for balancing convergence speed and computational cost. However, existing works typically choose $K$ empirically, without theoretical guidance. To address this gap, we characterize…

Cited by 0SourceScholar