← Search

Ioana Baldini

6 accepted papers

2025

SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models

NAACL 2025long

Large Language Models (LLMs) reproduce and exacerbate the social biases present in their training data, and resources to quantify this issue are limited. While research has attempted to identify and mitigate such biases, most efforts have been concentrated around English, lagging the rapid advanceme…

Cited by 1SourcePDFScholar
2024

Biasly: An Expert-Annotated Dataset for Subtle Misogyny Detection and Mitigation

ACL 2024findings

Using novel approaches to dataset development, the Biasly dataset captures the nuance and subtlety of misogyny in ways that are unique within the literature. Built in collaboration with multi-disciplinary experts and annotators themselves, the dataset contains annotations of movie subtitles, capturi…

2024

Fairness-Aware Structured Pruning in Transformers

AAAI 2024technical

The increasing size of large language models (LLMs) has introduced challenges in their training and inference. Removing model components is perceived as a solution to tackle the large model sizes, however, existing pruning methods solely focus on performance, without considering an essential aspect…

2024

SocialStigmaQA: A Benchmark to Uncover Stigma Amplification in Generative Language Models

AAAI 2024technical

Current datasets for unwanted social bias auditing are limited to studying protected demographic features such as race and gender. In this work, we introduce a comprehensive benchmark that is meant to capture the amplification of social bias, via stigmas, in generative language models. Taking inspir…

Cited by 18SourcePDFScholar
2024

Why Don’t Prompt-Based Fairness Metrics Correlate?

ACL 2024long

The widespread use of large language models has brought up essential questions about the potential biases these models might learn. This led to the development of several metrics aimed at evaluating and mitigating these biases. In this paper, we first demonstrate that prompt-based fairness metrics e…

2022

Your fairness may vary: Pretrained language model fairness in toxic text classification

ACL 2022findings

The popularity of pretrained language models in natural language processing systems calls for a careful evaluation of such models in down-stream tasks, which have a higher potential for societal impact. The evaluation of such systems usually focuses on accuracy measures. Our findings in this paper c…

Cited by 72SourcePDFScholar