2024
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
NeurIPS 2024poster
Larger transformer models perform better on various downstream tasks but require more cost to scale up the model size. To efficiently enlarge models, the Mixture-of-Expert (MoE) architecture is widely adopted, which consists of a gate network and a series of experts and keep the training cost consta…