← Search

James A. Michaelov

4 accepted papers

2025

Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale

NeurIPS 2025poster

We show that across architecture (Transformer vs. Mamba vs. RWKV), training dataset (OpenWebText vs. The Pile), and scale (14 million parameters to 12 billion parameters), autoregressive language models exhibit highly consistent patterns of change in their behavior over the course of pretraining. Ba…

Cited by 0SourceScholar
2025

Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events

ACL 2025finding

Can language models reliably predict that possible events are more likely than merely improbable ones? By teasing apart possibility, typicality, and contextual relatedness, we show that despite the results of previous work, language models’ ability to do this is far from robust. In fact, under certa…

Cited by 0SourcePDFScholar
2025

On the Acquisition of Shared Grammatical Representations in Bilingual Language Models

ACL 2025long

Crosslingual transfer is crucial to contemporary language models’ multilingual capabilities, but how it occurs is not well understood. Weask what happens to a monolingual language model when it begins to be trained on a second language. Specifically, we train small bilingual models for which we cont…

Cited by 0SourcePDFScholar
2022

Do Language Models Make Human-like Predictions about the Coreferents of Italian Anaphoric Zero Pronouns?

COLING 2022main

Some languages allow arguments to be omitted in certain contexts. Yet human language comprehenders reliably infer the intended referents of these zero pronouns, in part because they construct expectations about which referents are more likely. We ask whether Neural Language Models also extract the s…