2026
RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
ICML 2026poster
Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active when relevant to the task. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, i…