← Search

Vani Kanjirangat

2 accepted papers

2025

Tokenization and Representation Biases in Multilingual Models on Dialectal NLP Tasks

EMNLP 2025

Dialectal data are characterized by linguistic variation that appears small to humans but has a significant impact on the performance of models. This dialect gap has been related to various factors (e.g., data size, economic and social factors) whose impact, however, turns out to be inconsistent. In