← Search

Dan Vann

5 accepted papers

2025

Anecdoctoring: Automated Red-Teaming Across Language and Place

EMNLP 2025

Disinformation is among the top risks of generative artificial intelligence (AI) misuse. Global adoption of generative AI necessitates red-teaming evaluations (i.e., systematic adversarial probing) that are robust across diverse languages and cultures, but red-teaming datasets are commonly US- and E

Cited by 0SourcePDFScholar
2025

Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming

NeurIPS 2025poster

In this position paper we argue that conclusions drawn about relative system safety or attack method efficacy via AI red teaming are often not supported by evidence provided by attack success rate (ASR) comparisons. We show, through conceptual, theoretical, and empirical contributions, that many c…

Cited by 0SourceScholar
2025

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

ICML 2025poster

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons" (Roose, 2024). In this position paper, we argue that the ML community would benefit from l…

Cited by 0SourcePDFScholar
2025

Taxonomizing Representational Harms using Speech Act Theory

ACL 2025finding

Representational harms are widely recognized among fairness-related harms caused by generative language systems. However, their definitions are commonly under-specified. We make a theoretical contribution to the specification of representational harms by introducing a framework, grounded in speech a…

Cited by 0SourcePDFScholar
2023

FairPrism: Evaluating Fairness-Related Harms in Text Generation

ACL 2023long

It is critical to measure and mitigate fairness-related harms caused by AI text generation systems, including stereotyping and demeaning harms. To that end, we introduce FairPrism, a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering a diverse set of harm…