2023
Adaptive Gating in Mixture-of-Experts based Language Models
EMNLP 2023long main
Large language models have demonstrated exceptional language understanding capabilities in many NLP tasks. Sparsely activated mixture-of-experts (MoE) has emerged as a promising solution for scaling models while maintaining a constant number of computational operations. Existing MoE models adopt a f…