← Search

Sharon Levy

14 accepted papers

2025

LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education

NAACL 2025findings

With the increasing adoption of large language models (LLMs) in education, concerns about inherent biases in these models have gained prominence. We evaluate LLMs for bias in the personalized educational setting, specifically focusing on the models’ roles as “teachers.” We reveal significant biases…

2025

Making FETCH! Happen: Finding Emergent Dog Whistles Through Common Habitats

ACL 2025long

Dog whistles are coded expressions with dual meanings: one intended for the general public (outgroup) and another that conveys a specific message to an intended audience (ingroup). Often, these expressions are used to convey controversial political opinions while maintaining plausible deniability an…

2024

Evaluating Biases in Context-Dependent Sexual and Reproductive Health Questions

EMNLP 2024finding

Chat-based large language models have the opportunity to empower individuals lacking high-quality healthcare access to receive personalized information across a variety of topics. However, users may ask underspecified questions that require additional context for a model to correctly answer. We stud…

Cited by 2SourcePDFScholar
2024

Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts

EMNLP 2024finding

Large language models (LLMs) acquire beliefs about gender from training data and can therefore generate text with stereotypical gender attitudes. Prior studies have demonstrated model generations favor one gender or exhibit stereotypes about gender, but have not investigated the complex dynamics tha…

2024

Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts

NAACL 2024short

Benchmarks of the multilingual capabilities of text-to-image (T2I) models compare generated images prompted in a test language to an expected image distribution over a concept set. One such benchmark, “Conceptual Coverage Across Languages” (CoCo-CroLa), assesses the tangible noun inventory of T2I mo…

Cited by 3SourcePDFScholar
2023

ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models

EMNLP 2023long findings

As large language models are integrated into society, robustness toward a suite of prompts is increasingly important to maintain reliability in a high-variance environment.Robustness evaluations must comprehensively encapsulate the various settings in which a user may invoke an intelligent system. T…

Cited by 0SourcecodeScholar
2023

Comparing Biases and the Impact of Multilingual Training across Multiple Languages

EMNLP 2023long main

Studies in bias and fairness in natural language processing have primarily examined social biases within a single language and/or across few attributes (e.g. gender, race). However, biases can manifest differently across various languages for individual attributes. As a result, it is critical to exa…

Cited by 0SourceScholar
2023

Foveate, Attribute, and Rationalize: Towards Physically Safe and Trustworthy AI

ACL 2023findings

Users’ physical safety is an increasing concern as the market for intelligent systems continues to grow, where unconstrained systems may recommend users dangerous actions that can lead to serious injury. Covertly unsafe text is an area of particular interest, as such text may arise from everyday sce…

2023

WikiWhy: Answering and Explaining Cause-and-Effect Questions

ICLR 2023top-5%

As large language models (LLMs) grow larger and more sophisticated, assessing their "reasoning" capabilities in natural language grows more challenging. Recent question answering (QA) benchmarks that attempt to assess reasoning are often limited by a narrow scope of covered situations and subject ma…

Cited by 20SourcePDFScholar
2022

HybriDialogue: An Information-Seeking Dialogue Dataset Grounded on Tabular and Textual Data

ACL 2022findings

A pressing challenge in current dialogue systems is to successfully converse with users on topics with information distributed across different modalities. Previous work in multiturn dialogue systems has primarily focused on either text or table information. In more realistic scenarios, having a joi…

Cited by 26SourcePDFScholar
2022

Mitigating Covertly Unsafe Text within Natural Language Systems

EMNLP 2022finding

An increasingly prevalent problem for intelligent technologies is text safety, as uncontrolled systems may generate recommendations to their users that lead to injury or life-threatening consequences. However, the degree of explicitness of a generated statement that can cause physical harm varies. I…

Cited by 7SourcePDFScholar
2022

SafeText: A Benchmark for Exploring Physical Safety in Language Models

EMNLP 2022main

Understanding what constitutes safe text is an important issue in natural language processing and can often prevent the deployment of models deemed harmful and unsafe. One such type of safety that has been scarcely studied is commonsense physical safety, i.e. text that is not explicitly violent and…

2021

Modeling Disclosive Transparency in NLP Application Descriptions

EMNLP 2021main

Broader disclosive transparency—truth and clarity in communication regarding the function of AI systems—is widely considered desirable. Unfortunately, it is a nebulous concept, difficult to both define and quantify. This is problematic, as previous work has demonstrated possible trade-offs and negat…

2021

Open-Domain Question-Answering for COVID-19 and Other Emergent Domains

EMNLP 2021system demonstrations

Since late 2019, COVID-19 has quickly emerged as the newest biomedical domain, resulting in a surge of new information. As with other emergent domains, the discussion surrounding the topic has been rapidly changing, leading to the spread of misinformation. This has created the need for a public spac…