2024
XMoE: Sparse Models with Fine-grained and Adaptive Expert Selection
ACL 2024findings
Sparse models, including sparse Mixture-of-Experts (MoE) models, have emerged as an effective approach for scaling Transformer models. However, they often suffer from computational inefficiency since a significant number of parameters are unnecessarily involved in computations by multiplying values…