2024
How Far Is Too Far? Studying the Effects of Domain Discrepancy on Masked Language Models
COLING 2024main
Pre-trained masked language models, such as BERT, perform strongly on a wide variety of NLP tasks and have become ubiquitous in recent years. The typical way to use such models is to fine-tune them on downstream data. In this work, we aim to study how the difference in domains between the pre-traine…