2024
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
EMNLP 2024main
Language models strongly rely on frequency information because they maximize the likelihood of tokens during pre-training. As a consequence, language models tend to not generalize well to tokens that are seldom seen during training. Moreover, maximum likelihood training has been discovered to give r…