← Search

Tanja Samardzic

5 accepted papers

2025

Tokenization and Representation Biases in Multilingual Models on Dialectal NLP Tasks

EMNLP 2025

Dialectal data are characterized by linguistic variation that appears small to humans but has a significant impact on the performance of models. This dialect gap has been related to various factors (e.g., data size, economic and social factors) whose impact, however, turns out to be inconsistent. In

2024

A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets

NAACL 2024findings

Typologically diverse benchmarks are increasingly created to track the progress achieved in multilingual NLP. Linguistic diversity of these data sets is typically measured as the number of languages or language families included in the sample, but such measures do not consider structural properties…

2022

Subword Evenness (SuE) as a Predictor of Cross-lingual Transfer to Low-resource Languages

EMNLP 2022main

Pre-trained multilingual models, such as mBERT, XLM-R and mT5, are used to improve the performance on various tasks in low-resource languages via cross-lingual transfer. In this framework, English is usually seen as the most natural choice for a transfer language (for fine-tuning or continued traini…