← Search

Alexandra Olteanu

13 accepted papers

2025

Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems

ACL 2025long

As text generation systems’ outputs are increasingly anthropomorphic—perceived as human-like—scholars have also increasingly raised concerns about how such outputs can lead to harmful outcomes, such as users over-relying or developing emotional dependence on these systems. How to intervene on such s…

Cited by 0SourcePDFScholar
2025

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

ICML 2025poster

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons" (Roose, 2024). In this position paper, we argue that the ML community would benefit from l…

Cited by 0SourcePDFScholar
2025

Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor

NeurIPS 2025poster

In AI research and practice, rigor remains largely understood in terms of methodological rigor---such as whether mathematical, statistical, or computational methods are correctly applied. We argue that this narrow conception of rigor has contributed to the concerns raised by the responsible AI commu…

Cited by 0SourceScholar
2025

Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems

ACL 2025finding

The NLP research community has made publicly available numerous instruments for measuring representational harms caused by large language model (LLM)-based systems. These instruments have taken the form of datasets, metrics, tools, and more. In this paper, we examine the extent to which such instrum…

Cited by 0SourcePDFScholar
2024

Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective

ACL 2024findings

It is increasingly common to evaluate the same coreference resolution (CR) model on multiple datasets. Do these multi-dataset evaluations allow us to draw meaningful conclusions about model generalization? Or, do they rather reflect the idiosyncrasies of a particular experimental setup (e.g., the sp…

2024

ECBD: Evidence-Centered Benchmark Design for NLP

ACL 2024long

Benchmarking is seen as critical to assessing progress in NLP. However, creating a benchmark involves many design decisions (e.g., which datasets to include, which metrics to use) that often rely on tacit, untested assumptions about what the benchmark is intended to measure or is actually measuring.…

2024

“One-Size-Fits-All”? Examining Expectations around What Constitute “Fair” or “Good” NLG System Behaviors

NAACL 2024long

Fairness-related assumptions about what constitute appropriate NLG system behaviors range from invariance, where systems are expected to behave identically for social groups, to adaptation, where behaviors should instead vary across them. To illuminate tensions around invariance and adaptation, we c…

Cited by 7SourcePDFScholar
2023

FairPrism: Evaluating Fairness-Related Harms in Text Generation

ACL 2023long

It is critical to measure and mitigate fairness-related harms caused by AI text generation systems, including stereotyping and demeaning harms. To that end, we introduce FairPrism, a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering a diverse set of harm…

2023

Responsible AI Considerations in Text Summarization Research: A Review of Current Practices

EMNLP 2023long findings

AI and NLP publication venues have increasingly encouraged researchers to reflect on possible ethical considerations, adverse impacts, and other responsible AI issues their work might engender. However, for specific NLP tasks our understanding of how prevalent such issues are, or when and why these…

Cited by 0SourceScholar
2023

The KITMUS Test: Evaluating Knowledge Integration from Multiple Sources

ACL 2023long

Many state-of-the-art natural language understanding (NLU) models are based on pretrained neural language models. These models often make inferences using information from multiple sources. An important class of such inferences are those that require both background knowledge, presumably contained i…

2022

Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications

NAACL 2022long

There are many ways to express similar things in text, which makes evaluating natural language generation (NLG) systems difficult. Compounding this difficulty is the need to assess varying quality criteria depending on the deployment setting. While the landscape of NLG evaluation has been well-mappe…

Cited by 36SourcePDFScholar
2021

ADEPT: An Adjective-Dependent Plausibility Task

ACL 2021long

A false contract is more likely to be rejected than a contract is, yet a false key is less likely than a key to open doors. While correctly interpreting and assessing the effects of such adjective-noun pairs (e.g., false key) on the plausibility of given events (e.g., opening doors) underpins many n…

2021

Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets

ACL 2021long

Auditing NLP systems for computational harms like surfacing stereotypes is an elusive goal. Several recent efforts have focused on benchmark datasets consisting of pairs of contrastive sentences, which are often accompanied by metrics that aggregate an NLP system’s behavior on these pairs into measu…

Cited by 335SourcePDFScholar