NeurIPS 2020poster118 citations

Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping

Minjia Zhang, Yuxiong He

Abstract

Recently, Transformer-based language models have demonstrated remarkable performance across many NLP domains. However, the unsupervised pre-training step of these models suffers from unbearable overall computational expenses. Current methods for accelerating the pre-training either rely on massive parallelism with advanced hardware or are not applicable to language models.

BibTeX
@inproceedings{NEURIPS2020_a1140a3d,
 author = {Zhang, Minjia and He, Yuxiong},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {H. Larochelle and M. Ranzato and R. Hadsell and M.F. Balcan and H. Lin},
 pages = {14011--14023},
 publisher = {Curran Associates, Inc.},
 title = {Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping},
 url = {https://proceedings.neurips.cc/paper_files/paper/2020/file/a1140a3d0df1c81e24ae954d935e8926-Paper.pdf},
 volume = {33},
 year = {2020}
}