2022
AutoBERT-Zero: Evolving BERT Backbone from Scratch
AAAI 2022technical
Transformer-based pre-trained language models like BERT and its variants have recently achieved promising performance in various natural language processing (NLP) tasks. However, the conventional paradigm constructs the backbone by purely stacking the manually designed global self-attention layers,…