← Search

Valentin Barriere

7 accepted papers

2025

Adapting Bias Evaluation to Domain Contexts using Generative Models

EMNLP 2025

Numerous datasets have been proposed to evaluate social bias in Natural Language Processing (NLP) systems. However, assessing bias within specific application domains remains challenging, as existing approaches often face limitations in scalability and fidelity across domains. In this work, we intro

Cited by 0SourcePDFScholar
2025

StandUp4AI: A New Multilingual Dataset for Humor Detection in Stand-up Comedy Videos

EMNLP 2025

Aiming towards improving current computational models of humor detection, we propose a new multimodal dataset of stand-up comedies, in seven languages: English, French, Spanish, Italian, Portuguese, Hungarian and Czech. Our dataset of more than 330 hours %, is at the time of writing the biggest avai

2024

A Study of Nationality Bias in Names and Perplexity using Off-the-Shelf Affect-related Tweet Classifiers

EMNLP 2024main

In this paper, we apply a method to quantify biases associated with named entities from various countries. We create counterfactual examples with small perturbations on target-domain data instead of relying on templates or specific datasets for bias detection. On widely used classifiers for subjecti…

2024

Are Text Classifiers Xenophobic? A Country-Oriented Bias Detection Method with Least Confounding Variables

COLING 2024main

Classical bias detection methods used in Machine Learning are themselves biased because of the different confounding variables implied in the assessment of the initial biases. First they are using templates that are syntactically simple and distant from the target data on which the model will deploy…

Cited by 1SourcePDFScholar
2024

The Touché23-ValueEval Dataset for Identifying Human Values behind Arguments

COLING 2024main

While human values play a crucial role in making arguments persuasive, we currently lack the necessary extensive datasets to develop methods for analyzing the values underlying these arguments on a large scale. To address this gap, we present the Touché23-ValueEval dataset, an expansion of the Webis…

2023

Deep Natural Language Feature Learning for Interpretable Prediction

EMNLP 2023long main

We propose a general method to break down a main complex task into a set of intermediary easier sub-tasks, which are formulated in natural language as binary questions related to the final target task. Our method allows for representing each example by a vector consisting of the answers to these que…

Cited by 0SourcecodeScholar
2020

Improving Sentiment Analysis over non-English Tweets using Multilingual Transformers and Automatic Translation for Data-Augmentation

COLING 2020main

Tweets are specific text data when compared to general text. Although sentiment analysis over tweets has become very popular in the last decade for English, it is still difficult to find huge annotated corpora for non-English languages. The recent rise of the transformer models in Natural Language P…

Cited by 64SourcePDFScholar