← Search

Qiang Su

2 accepted papers

2023

Adaptive Gating in Mixture-of-Experts based Language Models

EMNLP 2023long main

Large language models have demonstrated exceptional language understanding capabilities in many NLP tasks. Sparsely activated mixture-of-experts (MoE) has emerged as a promising solution for scaling models while maintaining a constant number of computational operations. Existing MoE models adopt a f…

Cited by 0SourceScholar