← Search

Aida Mostafazadeh Davani

7 accepted papers

2025

A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI Evaluations

EMNLP 2025

Societal stereotypes are at the center of a myriad of responsible AI interventions targeted at reducing the generation and propagation of potentially harmful outcomes. While these efforts are much needed, they tend to be fragmented and often address different parts of the issue without adopting a un

Cited by 0SourcePDFScholar
2025

Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models

NeurIPS 2025spotlight

Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralism in AI alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions…

Cited by 0SourceScholar
2024

D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation

EMNLP 2024main

While human annotations play a crucial role in language technologies, annotator subjectivity has long been overlooked in data collection. Recent studies that critically examine this issue are often focused on Western contexts, and solely document differences across age, gender, or racial groups. Con…

Cited by 7SourcePDFScholar
2024

GRASP: A Disagreement Analysis Framework to Assess Group Associations in Perspectives

NAACL 2024long

Human annotation plays a core role in machine learning — annotations for supervised models, safety guardrails for generative models, and human feedback for reinforcement learning, to cite a few avenues. However, the fact that many of these human annotations are inherently subjective is often overloo…

2023

Distinguishing Address vs. Reference Mentions of Personal Names in Text

ACL 2023findings

Detecting named entities in text has long been a core NLP task. However, not much work has gone into distinguishing whether an entity mention is addressing the entity vs. referring to the entity; e.g., John, would you turn the light off? vs. John turned the light off. While this distinction is marke…

Cited by 1SourcePDFScholar
2023

SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative Models

ACL 2023long

Stereotype benchmark datasets are crucial to detect and mitigate social stereotypes about groups of people in NLP models. However, existing datasets are limited in size and coverage, and are largely restricted to stereotypes prevalent in the Western society. This is especially problematic as languag…

2021

On Transferability of Bias Mitigation Effects in Language Model Fine-Tuning

NAACL 2021long

Fine-tuned language models have been shown to exhibit biases against protected groups in a host of modeling tasks such as text classification and coreference resolution. Previous works focus on detecting these biases, reducing bias in data representations, and using auxiliary training objectives to…