← Search

Valerio Basile

12 accepted papers

2025

PERSEVAL: A Framework for Perspectivist Classification Evaluation

EMNLP 2025

Data perspectivism goes beyond majority vote label aggregation by recognizing various perspectives as legitimate ground truths.However, current evaluation practices remain fragmented, making it difficult to compare perspectivist approaches and analyze their impact on different users and demographic

2025

WebNLG-IT: Construction of an aligned RDF-Italian corpus through Machine Translation techniques

ACL 2025finding

The main goal of this work is the creation of the Italian version of the WebNLG corpus through the application of Neural Machine Translation (NMT) and post-editing with hand-written rules. To achieve this goal, in a first step, several existing NMT models were analysed and compared in order to ident…

2024

Capturing Perspectives of Crowdsourced Annotators in Subjective Learning Tasks

NAACL 2024long

Supervised classification heavily depends on datasets annotated by humans. However, in subjective tasks such as toxicity classification, these annotations often exhibit low agreement among raters. Annotations have commonly been aggregated by employing methods like majority voting to determine a sing…

2024

I’m sure you’re a real scholar yourself: Exploring Ironic Content Generation by Large Language Models

EMNLP 2024finding

Generating ironic content is challenging: it requires a nuanced understanding of context and implicit references and balancing seriousness and playfulness. Moreover, irony is highly subjective and can depend on various factors, such as social, cultural, or generational aspects. This paper explores w…

Cited by 0SourcePDFScholar
2024

Label Augmentation for Zero-Shot Hierarchical Text Classification

ACL 2024long

Hierarchical Text Classification poses the difficult challenge of classifying documents into multiple labels organized in a hierarchy. The vast majority of works aimed to address this problem relies on supervised methods which are difficult to implement due to the scarcity of labeled data in many re…

2024

MultiPICo: Multilingual Perspectivist Irony Corpus

ACL 2024long

Recently, several scholars have contributed to the growth of a new theoretical framework in NLP called perspectivism. This approach aimsto leverage data annotated by different individuals to model diverse perspectives that affect their opinions on subjective phenomena such as irony. In this context,…

2024

QUEEREOTYPES: A Multi-Source Italian Corpus of Stereotypes towards LGBTQIA+ Community Members

COLING 2024main

The paper describes a dataset composed of two sub-corpora from two different sources in Italian. The QUEEREOTYPES corpus includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events. The texts were collected from Facebook and Twitter in 2018 and were annotated for the p…

2023

Confidence-based Ensembling of Perspective-aware Models

EMNLP 2023long main

Research in the field of NLP has recently focused on the variability that people show in selecting labels when performing an annotation task. Exploiting disagreements in annotations has been shown to offer advantages for accurate modelling and fair evaluation. In this paper, we propose a strongly pe…

Cited by 0SourceScholar
2023

EPIC: Multi-Perspective Annotation of a Corpus of Irony

ACL 2023long

We present EPIC (English Perspectivist Irony Corpus), the first annotated corpus for irony analysis based on the principles of data perspectivism. The corpus contains short conversations from social media in five regional varieties of English, and it is annotated by contributors from five countries…

2023

Toward a Perspectivist Turn in Ground Truthing for Predictive Computing

AAAI 2023technical

Most current Artificial Intelligence applications are based on supervised Machine Learning (ML), which ultimately grounds on data annotated by small teams of experts or large ensemble of volunteers. The annotation process is often performed in terms of a majority vote, however this has been proved t…

Cited by 172SourcePDFScholar
2020

Multilingual Irony Detection with Dependency Syntax and Neural Models

COLING 2020main

This paper presents an in-depth investigation of the effectiveness of dependency-based syntactic features on the irony detection task in a multilingual perspective (English, Spanish, French and Italian). It focuses on the contribution from syntactic knowledge, exploiting linguistic resources where s…

2017

Semantic web-mining and deep vision for lifelong object discovery

ICRA 2017poster

Autonomous robots that are to assist humans in their daily lives must recognize and understand the meaning of objects in their environment. However, the open nature of the world means robots must be able to learn and extend their knowledge about previously unknown objects on-line. In this work we in…

Cited by 27SourceScholar