← Search

Eve Fleisig

12 accepted papers

2026

PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm

ICLR 2026poster

Current AI safety frameworks, which often treat harmfulness as binary, lack the flexibility to handle borderline cases where humans meaningfully disagree. To build more pluralistic systems, it is essential to move beyond consensus and instead understand where and why disagreements arise. We introduc…

Cited by 0SourceScholar
2025

GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration

ACL 2025long

Language models are often miscalibrated, leading to confidently incorrect answers. We introduce GRACE, a benchmark for language model calibration that incorporates comparison with human calibration. GRACE consists of question-answer pairs, in which each question contains a series of clues that gradu…

Cited by 0SourcePDFScholar
2025

Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness

NAACL 2025long

Adversarial datasets should validate AI robustness by providing samples on which humans perform well, but models do not. However, as models evolve, datasets can become obsolete. Measuring whether a dataset remains adversarial is hindered by the lack of a standardized metric for measuring adversarial…

Cited by 0SourcePDFScholar
2024

Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree

EMNLP 2024main

When annotators disagree, predicting the labels given by individual annotators can capture nuances overlooked by traditional label aggregation. We introduce three approaches to predict individual annotator ratings on the toxicity of text by incorporating individual annotator-specific information: a…

Cited by 1SourcePDFScholar
2024

First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models

NAACL 2024long

Many NLP researchers are experiencing an existential crisis triggered by the astonishing success of ChatGPT and other systems based on large language models (LLMs). After such a disruptive change to our understanding of the field, what is left to do? Taking a historical lens, we look for guidance fr…

Cited by 17SourcePDFScholar
2024

Ghostbuster: Detecting Text Ghostwritten by Large Language Models

NAACL 2024long

We introduce Ghostbuster, a state-of-the-art system for detecting AI-generated text.Our method works by passing documents through a series of weaker language models, running a structured search over possible combinations of their features, and then training a classifier on the selected features to p…

2024

Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination

EMNLP 2024main

We present a large-scale study of linguistic bias exhibited by ChatGPT covering ten dialects of English (Standard American English, Standard British English, and eight widely spoken non-”standard” varieties from around the world). We prompted GPT-3.5 Turbo and GPT-4 with text by native speakers of e…

Cited by 24SourcePDFScholar
2024

The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels

NAACL 2024long

Longstanding data labeling practices in machine learning involve collecting and aggregating labels from multiple annotators. But what should we do when annotators disagree? Though annotator disagreement has long been seen as a problem to minimize, new perspectivist approaches challenge this assumpti…

Cited by 19SourcePDFScholar
2023

Centering the Margins: Outlier-Based Identification of Harmed Populations in Toxicity Detection

EMNLP 2023long main

The impact of AI models on marginalized communities has traditionally been measured by identifying performance differences between specified demographic subgroups. Though this approach aims to center vulnerable groups, it risks obscuring patterns of harm faced by intersectional subgroups or shared a…

Cited by 0SourceScholar
2023

FairPrism: Evaluating Fairness-Related Harms in Text Generation

ACL 2023long

It is critical to measure and mitigate fairness-related harms caused by AI text generation systems, including stereotyping and demeaning harms. To that end, we introduce FairPrism, a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering a diverse set of harm…

2023

When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks

EMNLP 2023long main

Though majority vote among annotators is typically used for ground truth labels in machine learning, annotator disagreement in tasks such as hate speech detection may reflect systematic differences in opinion across groups, not noise. Thus, a crucial problem in hate speech detection is determining i…

Cited by 0SourceScholar