← Search

Ben Bergen

10 accepted papers

2025

Are explicit belief representations necessary? A comparison between Large Language Models and Bayesian probabilistic models

NAACL 2025long

Large language models (LLMs) have exhibited certain indirect pragmatic capabilities, including interpreting indirect requests and non-literal meanings. Yet, it is unclear whether the success of LLMs on pragmatic tasks generalizes to phenomena that directly probe inferences about the beliefs of other…

2025

Explaining and Mitigating Crosslingual Tokenizer Inequities

NeurIPS 2025poster

The number of tokens it takes to encode parallel text in different languages is known to vary. These disparities are called *token premiums*. Having high token premiums leads to less throughput during training and increases costs at inference. In this paper, we show that even after controlling for…

Cited by 0SourceScholar
2025

Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale

NeurIPS 2025poster

We show that across architecture (Transformer vs. Mamba vs. RWKV), training dataset (OpenWebText vs. The Pile), and scale (14 million parameters to 12 billion parameters), autoregressive language models exhibit highly consistent patterns of change in their behavior over the course of pretraining. Ba…

Cited by 0SourceScholar
2025

Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events

ACL 2025finding

Can language models reliably predict that possible events are more likely than merely improbable ones? By teasing apart possibility, typicality, and contextual relatedness, we show that despite the results of previous work, language models’ ability to do this is far from robust. In fact, under certa…

Cited by 0SourcePDFScholar
2025

On the Acquisition of Shared Grammatical Representations in Bilingual Language Models

ACL 2025long

Crosslingual transfer is crucial to contemporary language models’ multilingual capabilities, but how it occurs is not well understood. Weask what happens to a monolingual language model when it begins to be trained on a second language. Specifically, we train small bilingual models for which we cont…

Cited by 0SourcePDFScholar
2024

When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages

EMNLP 2024main

Multilingual language models are widely used to extend NLP systems to low-resource languages. However, concrete evidence for the effects of multilinguality on language modeling performance in individual languages remains scarce. Here, we pre-train over 10,000 monolingual and multilingual language mo…

2023

Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language Models

EMNLP 2023long main

Abstract grammatical knowledge—of parts of speech and grammatical patterns—is key to the capacity for linguistic generalization in humans. But how abstract is grammatical knowledge in large language models? In the human literature, compelling evidence for grammatical abstraction comes from structura…

Cited by 0SourceScholar