2025
HMoE: Heterogeneous Mixture of Experts for Language Modeling
EMNLP 2025
Mixture of Experts (MoE) offers remarkable performance and computational efficiency by selectively activating subsets of model parameters. Traditionally, MoE models use homogeneous experts, each with identical capacity. However, varying complexity in input data necessitates experts with diverse capa