Revisiting Efficiency–Accuracy Scaling in Mixture-of-Experts Architectures
Venmugil Elango, Nidhi Bhatia, Roger Waleffe, Rasoul Shafipour, Tomer Asida, Abhinav Khattar, Nave Assaf, Maximilian Golub
Abstract
Mixture-of-Experts (MoEs) have become a central component of many state-of-the-art open-source and proprietary large language models. Despite their widespread adoption, it remains unclear how close existing MoE architectures are to optimal with respect to inference cost, as measured by accuracy per floating-point operation and per parameter. In this work, we revisit MoE design from a hardware-software co-design perspective, grounded in empirical and theoretical considerations. We characterize key performance bottlenecks across diverse deployment regimes, spanning offline high-throughput execution and online, latency-critical inference. Guided by these insights, we introduce LatentMoE, a new model architecture resulting from systematic design exploration and optimized for maximized accuracy per unit of compute. Empirical design space exploration at scales of up to 95B parameters and over a 1T-token training horizon, together with supporting theoretical analysis, show that LatentMoE consistently outperforms standard MoE architectures in terms of accuracy per FLOP and per parameter.
BibTeX
@inproceedings{
elango2026revisiting,
title={Revisiting Efficiency{\textendash}Accuracy Scaling in Mixture-of-Experts Architectures},
author={Venmugil Elango and Nidhi Bhatia and Roger Waleffe and Rasoul Shafipour and Tomer Asida and Abhinav Khattar and Nave Assaf and Maximilian Golub and Joseph Guman and Tiyasa Mitra and Ritchie Zhao and Ritika Borkar and Ran Zilberstein and Mostofa Patwary and Mohammad Shoeybi and Bita Darvish Rouhani},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=iQRy0io4sb}
}