← Search

Tyler A. Chang

8 accepted papers

2025

Explaining and Mitigating Crosslingual Tokenizer Inequities

NeurIPS 2025poster

The number of tokens it takes to encode parallel text in different languages is known to vary. These disparities are called *token premiums*. Having high token premiums leads to less throughput during training and increases costs at inference. In this paper, we show that even after controlling for…

Cited by 0SourceScholar
2025

On the Acquisition of Shared Grammatical Representations in Bilingual Language Models

ACL 2025long

Crosslingual transfer is crucial to contemporary language models’ multilingual capabilities, but how it occurs is not well understood. Weask what happens to a monolingual language model when it begins to be trained on a second language. Specifically, we train small bilingual models for which we cont…

Cited by 0SourcePDFScholar
2025

Scalable Influence and Fact Tracing for Large Language Model Pretraining

ICLR 2025poster

Training data attribution (TDA) methods aim to attribute model outputs back to specific training examples, and the application of these methods to large language model (LLM) outputs could significantly advance model transparency and data curation. However, it has been challenging to date to apply th…

2024

Correlations between Multilingual Language Model Geometry and Crosslingual Transfer Performance

COLING 2024main

A common approach to interpreting multilingual language models is to evaluate their internal representations. For example, studies have found that languages occupy distinct subspaces in the models’ representation spaces, and geometric distances between languages often reflect linguistic properties s…

Cited by 0SourcePDFScholar
2024

Detecting Hallucination and Coverage Errors in Retrieval Augmented Generation for Controversial Topics

COLING 2024main

We explore a strategy to handle controversial topics in LLM-based chatbots based on Wikipedia’s Neutral Point of View (NPOV) principle: acknowledge the absence of a single true answer and surface multiple perspectives. We frame this as retrieval augmented generation, where perspectives are retrieved…

Cited by 12SourcePDFScholar
2024

When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages

EMNLP 2024main

Multilingual language models are widely used to extend NLP systems to low-resource languages. However, concrete evidence for the effects of multilinguality on language modeling performance in individual languages remains scarce. Here, we pre-train over 10,000 monolingual and multilingual language mo…

2023

Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language Models

EMNLP 2023long main

Abstract grammatical knowledge—of parts of speech and grammatical patterns—is key to the capacity for linguistic generalization in humans. But how abstract is grammatical knowledge in large language models? In the human literature, compelling evidence for grammatical abstraction comes from structura…

Cited by 0SourceScholar