← Search

Willem Zuidema

7 accepted papers

2025

PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs

ICLR 2025poster

The stability of language model pre-training and its effects on downstream performance are still understudied. Prior work shows that the training process can yield significantly different results in response to slight variations in initial conditions, e.g., the random seed. Crucially, the research c…

2024

DecoderLens: Layerwise Interpretation of Encoder-Decoder Transformers

NAACL 2024findings

In recent years, several interpretability methods have been proposed to interpret the inner workings of Transformer models at different levels of precision and complexity.In this work, we propose a simple but effective technique to analyze encoder-decoder Transformers. Our method, which we name Deco…

2024

Do Language Models Exhibit Human-like Structural Priming Effects?

ACL 2024findings

We explore which linguistic factors—at the sentence and token level—play an important role in influencing language model predictions, and investigate whether these are reflective of results found in humans and human corpora (Gries and Kootstra, 2017). We make use of the structural priming paradigm—w…

2023

Homophone Disambiguation Reveals Patterns of Context Mixing in Speech Transformers

EMNLP 2023long main

Transformers have become a key architecture in speech processing, but our understanding of how they build up representations of acoustic and linguistic structure is limited. In this study, we address this gap by investigating how measures of 'context-mixing' developed for text models can be adapted…

Cited by 0SourcecodeScholar
2023

Transparency at the Source: Evaluating and Interpreting Language Models With Access to the True Distribution

EMNLP 2023long findings

We present a setup for training, evaluating and interpreting neural language models, that uses artificial, language-like data. The data is generated using a massive probabilistic grammar (based on state-split PCFGs), that is itself derived from a large natural language corpus, but also provides us c…

Cited by 0SourcecodeScholar