← Search

Moin Nadeem

2 accepted papers

2023

MosaicBERT: A Bidirectional Encoder Optimized for Fast Pretraining

NeurIPS 2023poster

Although BERT-style encoder models are heavily used in NLP research, many researchers do not pretrain their own BERTs from scratch due to the high cost of training. In the past half-decade since BERT first rose to prominence, many advances have been made with other transformer architectures and trai…

2021

StereoSet: Measuring stereotypical bias in pretrained language models

ACL 2021long

A stereotype is an over-generalized belief about a particular group of people, e.g., Asians are good at math or African Americans are athletic. Such beliefs (biases) are known to hurt target groups. Since pretrained language models are trained on large real-world data, they are known to capture ster…