← Search

Mark Diaz

5 accepted papers

2026

Decoding Safety Feedback from Diverse Raters: A Data-driven Lens on Responsiveness to Severity

ICML 2026poster

Ensuring the safety of Generative AI requires a nuanced understanding of pluralistic viewpoints. In this paper, we introduce a novel data-driven approach for analyzing ordinal safety ratings in pluralistic settings. Specifically, we address the challenge of interpreting nuanced differences in safety…

Cited by 0SourceScholar
2025

Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models

NeurIPS 2025spotlight

Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralism in AI alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions…

Cited by 0SourceScholar
2024

D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation

EMNLP 2024main

While human annotations play a crucial role in language technologies, annotator subjectivity has long been overlooked in data collection. Recent studies that critically examine this issue are often focused on Western contexts, and solely document differences across age, gender, or racial groups. Con…

Cited by 7SourcePDFScholar
2024

GRASP: A Disagreement Analysis Framework to Assess Group Associations in Perspectives

NAACL 2024long

Human annotation plays a core role in machine learning — annotations for supervised models, safety guardrails for generative models, and human feedback for reinforcement learning, to cite a few avenues. However, the fact that many of these human annotations are inherently subjective is often overloo…

2024

STAR: SocioTechnical Approach to Red Teaming Language Models

EMNLP 2024main

This research introduces STAR, a sociotechnical framework that improves on current best practices for red teaming safety of large language models. STAR makes two key contributions: it enhances steerability by generating parameterised instructions for human red teamers, leading to improved coverage o…

Cited by 13SourcePDFScholar