← Search

Alan Ramponi

8 accepted papers

2025

Multilingual vs Crosslingual Retrieval of Fact-Checked Claims: A Tale of Two Approaches

EMNLP 2025

Retrieval of previously fact-checked claims is a well-established task, whose automation can assist professional fact-checkers in the initial steps of information verification. Previous works have mostly tackled the task monolingually, i.e., having both the input and the retrieved claims in the same

2025

Translation in the Hands of Many: Centering Lay Users in Machine Translation Interactions

EMNLP 2025

Converging societal and technical factors have transformed language technologies into user-facing applications used by the general public across languages. Machine Translation (MT) has become a global tool, with cross-lingual services now also supported by dialogue systems powered by multilingual La

Cited by 0SourcePDFScholar
2024

Delving into Qualitative Implications of Synthetic Data for Hate Speech Detection

EMNLP 2024main

The use of synthetic data for training models for a variety of NLP tasks is now widespread. However, previous work reports mixed results with regards to its effectiveness on highly subjective tasks such as hate speech detection. In this paper, we present an in-depth qualitative analysis of the poten…

2024

Variationist: Exploring Multifaceted Variation and Bias in Written Language Data

ACL 2024system demonstrations

Exploring and understanding language data is a fundamental stage in all areas dealing with human language. It allows NLP practitioners to uncover quality concerns and harmful biases in data before training, and helps linguists and social scientists to gain insight into language use and human behavio…

2022

Features or Spurious Artifacts? Data-centric Baselines for Fair and Robust Hate Speech Detection

NAACL 2022long

Avoiding to rely on dataset artifacts to predict hate speech is at the cornerstone of robust and fair hate speech detection. In this paper we critically analyze lexical biases in hate speech detection via a cross-platform study, disentangling various types of spurious and authentic artifacts and ana…

2021

From Masked Language Modeling to Translation: Non-English Auxiliary Tasks Improve Zero-shot Spoken Language Understanding

NAACL 2021long

The lack of publicly available evaluation data for low-resource languages limits progress in Spoken Language Understanding (SLU). As key tasks like intent classification and slot filling require abundant training data, it is desirable to reuse existing data in high-resource languages to develop mode…