← Search

Richard Diehl Martinez

2 accepted papers

2024

Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing

EMNLP 2024main

Language models strongly rely on frequency information because they maximize the likelihood of tokens during pre-training. As a consequence, language models tend to not generalize well to tokens that are seldom seen during training. Moreover, maximum likelihood training has been discovered to give r…

Cited by 1SourcePDFScholar
2024

Tending Towards Stability: Convergence Challenges in Small Language Models

EMNLP 2024finding

Increasing the number of parameters in language models is a common strategy to enhance their performance. However, smaller language models remain valuable due to their lower operational costs. Despite their advantages, smaller models frequently underperform compared to their larger counterparts, eve…