← Search

Daya Khudia

1 accepted papers

2023

MosaicBERT: A Bidirectional Encoder Optimized for Fast Pretraining

NeurIPS 2023poster

Although BERT-style encoder models are heavily used in NLP research, many researchers do not pretrain their own BERTs from scratch due to the high cost of training. In the past half-decade since BERT first rose to prominence, many advances have been made with other transformer architectures and trai…