2025
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
ICML 2025poster
Mixture of Experts (MoE) architectures have significantly increased computational efficiency in both research and real-world applications of large-scale machine learning models. However, their scalability and efficiency under memory constraints remain relatively underexplored. In this work, we prese…