2025
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
NeurIPS 2025spotlight
Mixture-of-Experts (MoE) architecture offers enhanced efficiency for Large Language Models (LLMs) with modularized computation, yet its inherent sparsity poses significant hardware deployment challenges, including memory locality issues, communication overhead, and inefficient computing resource uti…