ICML 2022spotlight22 citations

Gating Dropout: Communication-efficient Regularization for Sparsely Activated Transformers

Rui Liu, Young Jin Kim, Alexandre Muzio, Hany Hassan

Abstract

Sparsely activated transformers, such as Mixture of Experts (MoE), have received great interest due to their outrageous scaling capability which enables dramatical increases in model size without significant increases in computational cost. To achieve this, MoE models replace the feedforward sub-layer with Mixture-of-Experts sub-layer in transformers and use a gating network to route each token to its assigned experts. Since the common practice for efficient training of such models requires distributing experts and tokens across different machines, this routing strategy often incurs huge cross-machine communication cost because tokens and their assigned experts likely reside in different machines. In this paper, we propose

BibTeX
@InProceedings{pmlr-v162-liu22g,
  title = 	 {Gating Dropout: Communication-efficient Regularization for Sparsely Activated Transformers},
  author =       {Liu, Rui and Kim, Young Jin and Muzio, Alexandre and Hassan, Hany},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {13782--13792},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/liu22g/liu22g.pdf},
  url = 	 {https://proceedings.mlr.press/v162/liu22g.html},
  abstract = 	 {Sparsely activated transformers, such as Mixture of Experts (MoE), have received great interest due to their outrageous scaling capability which enables dramatical increases in model size without significant increases in computational cost. To achieve this, MoE models replace the feedforward sub-layer with Mixture-of-Experts sub-layer in transformers and use a gating network to route each token to its assigned experts. Since the common practice for efficient training of such models requires distributing experts and tokens across different machines, this routing strategy often incurs huge cross-machine communication cost because tokens and their assigned experts likely reside in different machines. In this paper, we propose
Gating Dropout: Communication-efficient Regularization for Sparsely Activated Transformers · ICML 2022