NeurIPS 2020poster118 citations
Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping
Abstract
Recently, Transformer-based language models have demonstrated remarkable performance across many NLP domains. However, the unsupervised pre-training step of these models suffers from unbearable overall computational expenses. Current methods for accelerating the pre-training either rely on massive parallelism with advanced hardware or are not applicable to language models.
BibTeX
@inproceedings{NEURIPS2020_a1140a3d,
author = {Zhang, Minjia and He, Yuxiong},
booktitle = {Advances in Neural Information Processing Systems},
editor = {H. Larochelle and M. Ranzato and R. Hadsell and M.F. Balcan and H. Lin},
pages = {14011--14023},
publisher = {Curran Associates, Inc.},
title = {Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping},
url = {https://proceedings.neurips.cc/paper_files/paper/2020/file/a1140a3d0df1c81e24ae954d935e8926-Paper.pdf},
volume = {33},
year = {2020}
}