2025
FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
NeurIPS 2025poster
The parameter size of modern large language models (LLMs) can be scaled up to the trillion-level via the sparsely-activated Mixture-of-Experts (MoE) technique to avoid excessive increase of the computational costs. To further improve training efficiency, pipelining computation and communication has…