← Search

Fabio Rinaldi

3 accepted papers

2025

Tokenization and Representation Biases in Multilingual Models on Dialectal NLP Tasks

EMNLP 2025

Dialectal data are characterized by linguistic variation that appears small to humans but has a significant impact on the performance of models. This dialect gap has been related to various factors (e.g., data size, economic and social factors) whose impact, however, turns out to be inconsistent. In

2024

BUST: Benchmark for the evaluation of detectors of LLM-Generated Text

NAACL 2024long

We introduce BUST, a comprehensive benchmark designed to evaluate detectors of texts generated by instruction-tuned large language models (LLMs). Unlike previous benchmarks, our focus lies on evaluating the performance of detector systems, acknowledging the inevitable influence of the underlying tas…