← Search

Su Lin Blodgett

17 accepted papers

2025

Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems

ACL 2025long

As text generation systems’ outputs are increasingly anthropomorphic—perceived as human-like—scholars have also increasingly raised concerns about how such outputs can lead to harmful outcomes, such as users over-relying or developing emotional dependence on these systems. How to intervene on such s…

Cited by 0SourcePDFScholar
2025

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

ICML 2025poster

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons" (Roose, 2024). In this position paper, we argue that the ML community would benefit from l…

Cited by 0SourcePDFScholar
2025

Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor

NeurIPS 2025poster

In AI research and practice, rigor remains largely understood in terms of methodological rigor---such as whether mathematical, statistical, or computational methods are correctly applied. We argue that this narrow conception of rigor has contributed to the concerns raised by the responsible AI commu…

Cited by 0SourceScholar
2025

Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems

ACL 2025finding

The NLP research community has made publicly available numerous instruments for measuring representational harms caused by large language model (LLM)-based systems. These instruments have taken the form of datasets, metrics, tools, and more. In this paper, we examine the extent to which such instrum…

Cited by 0SourcePDFScholar
2024

ECBD: Evidence-Centered Benchmark Design for NLP

ACL 2024long

Benchmarking is seen as critical to assessing progress in NLP. However, creating a benchmark involves many design decisions (e.g., which datasets to include, which metrics to use) that often rely on tacit, untested assumptions about what the benchmark is intended to measure or is actually measuring.…

2024

Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLP

EMNLP 2024main

This paper introduces the concept of actionability in the context of bias measures in natural language processing (NLP). We define actionability as the degree to which a measure’s results enable informed action and propose a set of desiderata for assessing it. Building on existing frameworks such as…

Cited by 1SourcePDFScholar
2024

The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels

NAACL 2024long

Longstanding data labeling practices in machine learning involve collecting and aggregating labels from multiple annotators. But what should we do when annotators disagree? Though annotator disagreement has long been seen as a problem to minimize, new perspectivist approaches challenge this assumpti…

Cited by 19SourcePDFScholar
2024

Understanding the Impacts of Language Technologies’ Performance Disparities on African American Language Speakers

ACL 2024findings

This paper examines the experiences of African American Language (AAL) speakers when using language technologies. Previous work has used quantitative methods to uncover performance disparities between AAL speakers and White Mainstream English speakers when using language technologies, but has not so…

Cited by 12SourcePDFScholar
2024

“One-Size-Fits-All”? Examining Expectations around What Constitute “Fair” or “Good” NLG System Behaviors

NAACL 2024long

Fairness-related assumptions about what constitute appropriate NLG system behaviors range from invariance, where systems are expected to behave identically for social groups, to adaptation, where behaviors should instead vary across them. To illuminate tensions around invariance and adaptation, we c…

Cited by 7SourcePDFScholar
2023

FairPrism: Evaluating Fairness-Related Harms in Text Generation

ACL 2023long

It is critical to measure and mitigate fairness-related harms caused by AI text generation systems, including stereotyping and demeaning harms. To that end, we introduce FairPrism, a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering a diverse set of harm…

2023

It Takes Two to Tango: Navigating Conceptualizations of NLP Tasks and Measurements of Performance

ACL 2023findings

Progress in NLP is increasingly measured through benchmarks; hence, contextualizing progress requires understanding when and why practitioners may disagree about the validity of benchmarks. We develop a taxonomy of disagreement, drawing on tools from measurement modeling, and distinguish between two…

Cited by 16SourcePDFScholar
2023

Responsible AI Considerations in Text Summarization Research: A Review of Current Practices

EMNLP 2023long findings

AI and NLP publication venues have increasingly encouraged researchers to reflect on possible ethical considerations, adverse impacts, and other responsible AI issues their work might engender. However, for specific NLP tasks our understanding of how prevalent such issues are, or when and why these…

Cited by 0SourceScholar
2023

Taxonomizing and Measuring Representational Harms: A Look at Image Tagging

AAAI 2023technical

In this paper, we examine computational approaches for measuring the "fairness" of image tagging systems, finding that they cluster into five distinct categories, each with its own analytic foundation. We also identify a range of normative concerns that are often collapsed under the terms "unfairnes…

Cited by 52SourcePDFScholar
2023

This prompt is measuring <mask>: evaluating bias evaluation in language models

ACL 2023findings

Bias research in NLP seeks to analyse models for social biases, thus helping NLP practitioners uncover, measure, and mitigate social harms. We analyse the body of work that uses prompts and templates to assess bias in language models. We draw on a measurement modelling framework to create a taxonomy…

Cited by 34SourcePDFScholar
2022

Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications

NAACL 2022long

There are many ways to express similar things in text, which makes evaluating natural language generation (NLG) systems difficult. Compounding this difficulty is the need to assess varying quality criteria depending on the deployment setting. While the landscape of NLG evaluation has been well-mappe…

Cited by 36SourcePDFScholar
2021

Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets

ACL 2021long

Auditing NLP systems for computational harms like surfacing stereotypes is an elusive goal. Several recent efforts have focused on benchmark datasets consisting of pairs of contrastive sentences, which are often accompanied by metrics that aggregate an NLP system’s behavior on these pairs into measu…

Cited by 335SourcePDFScholar