← Search

Hanna Wallach

19 accepted papers

2025

Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming

NeurIPS 2025poster

In this position paper we argue that conclusions drawn about relative system safety or attack method efficacy via AI red teaming are often not supported by evidence provided by attack success rate (ASR) comparisons. We show, through conceptual, theoretical, and empirical contributions, that many c…

Cited by 0SourceScholar
2025

Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research

NeurIPS 2025oral

"Machine unlearning" is a popular proposed solution for mitigating the existence of content in an AI model that is problematic for legal or moral reasons, including privacy, copyright, safety, and more. For example, unlearning is often invoked as a solution for removing the effects of specific infor…

Cited by 0SourceScholar
2025

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

ICML 2025poster

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons" (Roose, 2024). In this position paper, we argue that the ML community would benefit from l…

Cited by 0SourcePDFScholar
2025

Taxonomizing Representational Harms using Speech Act Theory

ACL 2025finding

Representational harms are widely recognized among fairness-related harms caused by generative language systems. However, their definitions are commonly under-specified. We make a theoretical contribution to the specification of representational harms by introducing a framework, grounded in speech a…

Cited by 0SourcePDFScholar
2025

Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems

ACL 2025finding

The NLP research community has made publicly available numerous instruments for measuring representational harms caused by large language model (LLM)-based systems. These instruments have taken the form of datasets, metrics, tools, and more. In this paper, we examine the extent to which such instrum…

Cited by 0SourcePDFScholar
2025

Validating LLM-as-a-Judge Systems under Rating Indeterminacy

NeurIPS 2025poster

The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, plays a critical role in scaling and standardizing GenAI evaluations. To validate such judge systems, evaluators assess human--judge agreement by first collect…

Cited by 0SourceScholar
2024

Understanding the Impacts of Language Technologies’ Performance Disparities on African American Language Speakers

ACL 2024findings

This paper examines the experiences of African American Language (AAL) speakers when using language technologies. Previous work has used quantitative methods to uncover performance disparities between AAL speakers and White Mainstream English speakers when using language technologies, but has not so…

Cited by 12SourcePDFScholar
2024

“One-Size-Fits-All”? Examining Expectations around What Constitute “Fair” or “Good” NLG System Behaviors

NAACL 2024long

Fairness-related assumptions about what constitute appropriate NLG system behaviors range from invariance, where systems are expected to behave identically for social groups, to adaptation, where behaviors should instead vary across them. To illuminate tensions around invariance and adaptation, we c…

Cited by 7SourcePDFScholar
2023

FairPrism: Evaluating Fairness-Related Harms in Text Generation

ACL 2023long

It is critical to measure and mitigate fairness-related harms caused by AI text generation systems, including stereotyping and demeaning harms. To that end, we introduce FairPrism, a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering a diverse set of harm…

2023

Taxonomizing and Measuring Representational Harms: A Look at Image Tagging

AAAI 2023technical

In this paper, we examine computational approaches for measuring the "fairness" of image tagging systems, finding that they cluster into five distinct categories, each with its own analytic foundation. We also identify a range of normative concerns that are often collapsed under the terms "unfairnes…

Cited by 52SourcePDFScholar
2021

Doubly non-central beta matrix factorization for DNA methylation data

UAI 2021poster

We present a new non-negative matrix factorization model for $(0,1)$ bounded-support data based on the doubly non-central beta (DNCB) distribution, a generalization of the beta distribution. The expressiveness of the DNCB distribution is particularly useful for modeling DNA methylation datasets, whi…

2021

Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets

ACL 2021long

Auditing NLP systems for computational harms like surfacing stereotypes is an elusive goal. Several recent efforts have focused on benchmark datasets consisting of pairs of contrastive sentences, which are often accompanied by metrics that aggregate an NLP system’s behavior on these pairs into measu…

Cited by 335SourcePDFScholar
2019

Locally Private Bayesian Inference for Count Models

ICML 2019oral

We present a general and modular method for privacy-preserving Bayesian inference for Poisson factorization, a broad class of models that includes some of the most widely used models in the social sciences. Our method satisfies limited-precision local privacy, a generalization of local differential…

Cited by 43SourcePDFScholar
2019

Poisson-Randomized Gamma Dynamical Systems

NeurIPS 2019poster

This paper presents the Poisson-randomized gamma dynamical system (PRGDS), a model for sequentially observed count tensors that encodes a strong inductive bias toward sparsity and burstiness. The PRGDS is based on a new motif in Bayesian latent variable modeling, an alternating chain of discrete Poi…

2018

A Reductions Approach to Fair Classification

ICML 2018oral

We present a systematic approach for achieving fairness in a binary classification setting. While we focus on two well-known quantitative definitions of fairness, our approach encompasses many other previously studied definitions as special cases. The key idea is to reduce fair classification to a s…

2016

Bayesian Poisson Tucker Decomposition for Learning the Structure of International Relations

ICML 2016poster

We introduce Bayesian Poisson Tucker decomposition (BPTD) for modeling country–country interaction event data. These data consist of interaction events of the form “country i took action a toward country j at time t.” BPTD discovers overlapping country–community memberships, including the number of…

Cited by 105SourcePDFScholar
2016

Flexible Models for Microclustering with Application to Entity Resolution

NeurIPS 2016poster

Most generative models for clustering implicitly assume that the number of data points in each cluster grows linearly with the total number of data points. Finite mixture models, Dirichlet process mixture models, and Pitman--Yor process mixture models make this assumption, as do all other infinitely…

Cited by 60SourcePDFScholar
2015

The Bayesian Echo Chamber: Modeling Social Influence via Linguistic Accommodation

AISTATS 2015poster

We present the Bayesian Echo Chamber, a new Bayesian generative model for social interaction data. By modeling the evolution of people’s language usage over time, this model discovers latent influence relationships between them. Unlike previous work on inferring influence, which has primarily focuse…

Cited by 75SourcePDFScholar