← Search

Catherine Arnett

8 accepted papers

2026

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

ICLR 2026oral

Large Language Models (LLMs) are pre-trained on large data from different sources and domains. These data most often contain trillions of tokens with large portions of copyrighted or proprietary content, which hinders the usage of such models under AI legislation. This raises the need for truly open…

Cited by 0SourceScholar
2026

Position: Don't Just "Fix it in Post'': A Science of AI Must Study Learning Dynamics

ICML 2026oral

What would it mean to have a *scientific* understanding of AI? Language models are not static objects—they are snapshots of time-evolving processes shaped by data, objectives, and optimization dynamics. Yet the field predominantly treats models as fixed artifacts, analyzing behaviors after training …

Cited by 0SourceScholar
2025

Explaining and Mitigating Crosslingual Tokenizer Inequities

NeurIPS 2025poster

The number of tokens it takes to encode parallel text in different languages is known to vary. These disparities are called *token premiums*. Having high token premiums leads to less throughput during training and increases costs at inference. In this paper, we show that even after controlling for…

Cited by 0SourceScholar
2025

On the Acquisition of Shared Grammatical Representations in Bilingual Language Models

ACL 2025long

Crosslingual transfer is crucial to contemporary language models’ multilingual capabilities, but how it occurs is not well understood. Weask what happens to a monolingual language model when it begins to be trained on a second language. Specifically, we train small bilingual models for which we cont…

Cited by 0SourcePDFScholar
2025

Why do language models perform worse for morphologically complex languages?

COLING 2025main

Language models perform differently across languages. It has been previously suggested that morphological typology may explain some of this variability (Cotterell et al., 2018). We replicate previous analyses and find additional new evidence for a performance gap between agglutinative and fusional l…

2024

BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training

EMNLP 2024main

Language models can greatly benefit from efficient tokenization. However, they still mostly utilize the classical Byte-Pair Encoding (BPE) algorithm, a simple and reliable method. BPE has been shown to cause such issues as under-trained tokens and sub-optimal compression that may affect the downstre…

2024

When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages

EMNLP 2024main

Multilingual language models are widely used to extend NLP systems to low-resource languages. However, concrete evidence for the effects of multilinguality on language modeling performance in individual languages remains scarce. Here, we pre-train over 10,000 monolingual and multilingual language mo…

2023

Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language Models

EMNLP 2023long main

Abstract grammatical knowledge—of parts of speech and grammatical patterns—is key to the capacity for linguistic generalization in humans. But how abstract is grammatical knowledge in large language models? In the human literature, compelling evidence for grammatical abstraction comes from structura…

Cited by 0SourceScholar