← Search

Myra Cheng

7 accepted papers

2026

ELEPHANT: Measuring and understanding social sycophancy in LLMs

ICLR 2026poster

LLMs are known to exhibit _sycophancy_: agreeing with and flattering users, even at the cost of correctness. Prior work measures sycophancy only as direct agreement with users' explicitly stated beliefs that can be compared to a ground truth. This fails to capture broader forms of sycophancy such as…

Cited by 0SourcecodeScholar
2025

Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems

ACL 2025long

As text generation systems’ outputs are increasingly anthropomorphic—perceived as human-like—scholars have also increasingly raised concerns about how such outputs can lead to harmful outcomes, such as users over-relying or developing emotional dependence on these systems. How to intervene on such s…

Cited by 0SourcePDFScholar
2024

NLP Systems That Can’t Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps

NAACL 2024long

The use of words to convey speaker’s intent is traditionally distinguished from the ‘mention’ of words for quoting what someone said, or pointing out properties of a word. Here we show that computationally modeling this use-mention distinction is crucial for dealing with counterspeech online. Counte…

2023

CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations

EMNLP 2023long main

Recent work has aimed to capture nuances of human behavior by using LLMs to simulate responses from particular demographics in settings like social science experiments and public opinion surveys. However, there are currently no established ways to discuss or evaluate the quality of such LLM simulati…

Cited by 0SourcecodeScholar
2023

Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models

ACL 2023long

To recognize and mitigate harms from large language models (LLMs), we need to understand the prevalence and nuances of stereotypes in LLM outputs. Toward this end, we present Marked Personas, a prompt-based method to measure stereotypes in LLMs for intersectional demographic groups without any lexic…

2020

Human Preference-Based Learning for High-dimensional Optimization of Exoskeleton Walking Gaits

IROS 2020poster

Optimizing lower-body exoskeleton walking gaits for user comfort requires understanding users' preferences over a high-dimensional gait parameter space. However, existing preference-based learning methods have only explored low-dimensional domains due to computational limitations. To learn user pref…

Cited by 48SourcecodeScholar