ConLoan: A Contrastive Multilingual Dataset for Evaluating Loanwords
Sina Ahmadi, Micha David Hess, Elena Álvarez-Mellado, Alessia Battisti, Cui Ding, Anne Göhring, Yingqiang Gao, Zifan Jiang
Abstract
Lexical borrowing, the adoption of words from one language into another, is a ubiquitous linguistic phenomenon influenced by geopolitical, societal, and technological factors. This paper introduces ConLoan–a novel contrastive dataset comprising sentences with and without loanwords across 10 languages. Through systematic evaluation using this dataset, we investigate how state-of-the-art machine translation and language models process loanwords compared to their native alternatives. Our experiments reveal that these systems show systematic preferences for loanwords over native terms and exhibit varying performance across languages. These findings provide valuable insights for developing more linguistically robust NLP systems.
BibTeX
@inproceedings{ahmadi-etal-2025-conloan,
title = "{C}on{L}oan: A Contrastive Multilingual Dataset for Evaluating Loanwords",
author = {Ahmadi, Sina and
Hess, Micha David and
{\'A}lvarez-Mellado, Elena and
Battisti, Alessia and
Ding, Cui and
G{\"o}hring, Anne and
Gao, Yingqiang and
Jiang, Zifan and
Michail, Andrianos and
Morad, Peshmerge and
Niklaus, Joel and
Panagiotopoulou, Maria Christina and
Perrella, Stefano and
Opitz, Juri and
Shaitarova, Anastassia and
Sennrich, Rico},
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.acl-long.1453/",
doi = "10.18653/v1/2025.acl-long.1453",
pages = "30070--30090",
ISBN = "979-8-89176-251-0"
}