← Search

Katharina Hämmerl

8 accepted papers

2025

Beyond Literal Token Overlap: Token Alignability for Multilinguality

NAACL 2025short

Previous work has considered token overlap, or even similarity of token distributions, as predictors for multilinguality and cross-lingual knowledge transfer in language models. However, these very literal metrics assign large distances to language pairs with different scripts, which can nevertheles…

Cited by 0SourcePDFScholar
2025

Improving Parallel Sentence Mining for Low-Resource and Endangered Languages

ACL 2025short

While parallel sentence mining has been extensively covered for fairly well-resourced languages, pairs involving low-resource languages have received comparatively little attention.To address this gap, we present Belopsem, a benchmark of new datasets for parallel sentence mining on three language pa…

2025

Multilingual Text-to-Image Generation Magnifies Gender Stereotypes

ACL 2025long

Text-to-image (T2I) generation models have achieved great results in image quality, flexibility, and text alignment, leading to widespread use. Through improvements in multilingual abilities, a larger community can access this technology. Yet, we show that multilingual models suffer from substantial…

2023

A Study on Accessing Linguistic Information in Pre-Trained Language Models by Using Prompts

EMNLP 2023short main

We study whether linguistic information in pre-trained multilingual language models can be accessed by human language: So far, there is no easy method to directly obtain linguistic information and gain insights into the linguistic principles encoded in such models. We use the technique of prompting…

Cited by 0SourceScholar
2023

Exploring Anisotropy and Outliers in Multilingual Language Models for Cross-Lingual Semantic Sentence Similarity

ACL 2023findings

Previous work has shown that the representations output by contextual language models are more anisotropic than static type embeddings, and typically display outlier dimensions. This seems to be true for both monolingual and multilingual models, although much less work has been done on the multiling…

2023

Speaking Multiple Languages Affects the Moral Bias of Language Models

ACL 2023findings

Pre-trained multilingual language models (PMLMs) are commonly used when dealing with data from multiple languages and cross-lingual transfer. However, PMLMs are trained on varying amounts of data for each language. In practice this means their performance is often much better on English than many ot…

2022

Combining Static and Contextualised Multilingual Embeddings

ACL 2022findings

Static and contextual multilingual embeddings have complementary strengths. Static embeddings, while less expressive than contextual language models, can be more straightforwardly aligned across multiple languages. We combine the strengths of static and contextual models to improve multilingual repr…