← Search

Rachel Rudinger

29 accepted papers

2025

A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users

EMNLP 2025

To assist users in complex tasks, LLMs generate plans: step-by-step instructions towards a goal. While alignment methods aim to ensure LLM plans are helpful, they train (RLHF) or evaluate (ChatbotArena) on what users prefer, assuming this reflects what helps them. We test this with Planorama: an int

Cited by 0SourcePDFScholar
2025

Language Models Predict Empathy Gaps Between Social In-groups and Out-groups

NAACL 2025long

Studies of human psychology have demonstrated that people are more motivated to extend empathy to in-group members than out-group members (Cikara et al., 2011). In this study, we investigate how this aspect of intergroup relations in humans is replicated by LLMs in an emotion intensity prediction ta…

2025

Multiple LLM Agents Debate for Equitable Cultural Alignment

ACL 2025long

Large Language Models (LLMs) need to adapt their predictions to diverse cultural contexts to benefit diverse communities across the world. While previous efforts have focused on single-LLM, single-turn approaches, we propose to exploit the complementary strengths of multiple LLMs to promote cultural…

2025

Natural Language Inference Improves Compositionality in Vision-Language Models

ICLR 2025poster

Compositional reasoning in Vision-Language Models (VLMs) remains challenging as these models often struggle to relate objects, attributes, and spatial relationships. Recent methods aim to address these limitations by relying on the semantics of the textual description, using Large Language Models (L…

Cited by 3SourcePDFScholar
2025

No Questions are Stupid, but some are Poorly Posed: Understanding Poorly-Posed Information-Seeking Questions

ACL 2025long

Questions help unlock information to satisfy users’ information needs. However, when the question is poorly posed, answerers (whether human or computer) may struggle to answer the question in a way that satisfies the asker, despite possibly knowing everything necessary to address the asker’s latent…

2025

On the Mutual Influence of Gender and Occupation in LLM Representations

ACL 2025long

We examine LLM representations of gender for first names in various occupational contexts to study how occupations and the gender perception of first names in LLMs influence each other mutually. We find that LLMs’ first-name gender representations correlate with real-world gender statistics associat…

Cited by 0SourcePDFScholar
2025

Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can’t Answer?

NAACL 2025short

Question answering (QA)—giving correct answers to questions—is a popular task, but we test **reverse question answering (RQA)**: for an input answer, give a question with that answer. Past work tests QA and RQA separately, but we test them jointly, comparing their difficulty, aiding benchmark design…

2025

Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat Logs

ACL 2025long

While it is commonly accepted that maintaining common ground plays a role in conversational success, little prior research exists connecting conversational grounding to success in task-oriented conversations. We study failures of grounding in the Ubuntu IRC dataset, where participants use text-only…

Cited by 0SourcePDFScholar
2025

Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above

ACL 2025long

Multiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing, but we argue for its reform. We first reveal flaws in MCQA’s format, as it struggles to: 1) test generation/subjectivity; 2) match LLM use cases; and 3) fully test knowledge. We inst…

Cited by 0SourcePDFScholar
2025

Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas

ACL 2025long

LLMs are aligned to follow input instructions by learning which of two responses users prefer for a prompt. However, such preference data do not convey *why* users prefer responses that are chosen or rejected, so LLMs trained on these datasets cannot tailor responses to varied user needs. To surface…

2025

‘Rich Dad, Poor Lad’: How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?

EMNLP 2025

Large Language Models (LLMs) are increasingly involved in high-stakes domains, yet how they reason about socially-sensitive decisions still remain underexplored. We present a large-scale audit of LLMs’ treatment of socioeconomic status (SES) in college admissions decisions using a novel dual-process

Cited by 0SourcePDFScholar
2024

Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?

ACL 2024long

Multiple-choice question answering (MCQA) is often used to evaluate large language models (LLMs). To see if MCQA assesses LLMs as intended, we probe if LLMs can perform MCQA with choices-only prompts, where models must select the correct answer only from the choices. In three MCQA datasets and four…

2024

Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race, Ethnicity, and Gender?

ACL 2024short

We examine whether large language models (LLMs) exhibit race- and gender-based name discrimination in hiring decisions, similar to classic findings in the social sciences (Bertrand and Mullainathan, 2004). We design a series of templatic prompts to LLMs to write an email to a named job applicant inf…

2024

It’s Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning

ACL 2024findings

Chain-of-thought (COT) prompting can help large language models (LLMs) reason toward correct answers, but its efficacy in reasoning toward incorrect answers is unexplored. This process of elimination (PoE), when used with COT, can enhance self-consistency, interpretability, and tasks such as medical…

2024

On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models

EMNLP 2024main

We study the presence of heteronormative biases and prejudice against interracial romantic relationships in large language models by performing controlled name-replacement experiments for the task of relationship prediction. We show that models are less likely to predict romantic relationships for (…

2024

Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning

EMNLP 2024finding

Questions involving commonsense reasoning about everyday situations often admit many possible or plausible answers. In contrast, multiple-choice question (MCQ) benchmarks for commonsense reasoning require a hard selection of a single correct answer, which, in principle, should represent the most pla…

2024

Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering

NAACL 2024long

Questions posed by information-seeking users often contain implicit false or potentially harmful assumptions. In a high-risk domain such as maternal and infant health, a question-answering system must recognize these pragmatic constraints and go beyond simply answering user questions, examining them…

2024

Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the US

EMNLP 2024main

Recent work has highlighted the culturally-contingent nature of commonsense knowledge. We introduce AMAMMERε, a test set of 525 multiple-choice questions designed to evaluate the commonsense knowledge of English LLMs, relative to the cultural contexts of Ghana and the United States. To create AMAMME…

Cited by 0SourcePDFScholar
2023

FORK: A Bite-Sized Test Set for Probing Culinary Cultural Biases in Commonsense Reasoning Models

ACL 2023findings

It is common sense that one should prefer to eat a salad with a fork rather than with a chainsaw. However, for eating a bowl of rice, the choice between a fork and a pair of chopsticks is culturally relative. We introduce FORK, a small, manually-curated set of CommonsenseQA-style questions for probi…

Cited by 38SourcePDFScholar
2023

Nichelle and Nancy: The Influence of Demographic Attributes and Tokenization Length on First Name Biases

ACL 2023short

Through the use of first name substitution experiments, prior research has demonstrated the tendency of social commonsense reasoning models to systematically exhibit social biases along the dimensions of race, ethnicity, and gender (An et al., 2023). Demographic attributes of first names, however, a…

Cited by 12SourcePDFScholar
2023

What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and Prohibitions

EMNLP 2023long main

Reviewing and comprehending key obligations, entitlements, and prohibitions in legal contracts can be a tedious task due to their length and domain-specificity. Furthermore, the key rights and duties requiring review vary for each contracting party. In this work, we propose a new task of \textit{par…

Cited by 0SourceScholar
2022

Agent-Specific Deontic Modality Detection in Legal Language

EMNLP 2022main

Legal documents are typically long and written in legalese, which makes it particularly difficult for laypeople to understand their rights and duties. While natural language understanding technologies can be valuable in supporting such understanding in the legal domain, the limited availability of d…

2022

Partial-input baselines show that NLI models can ignore context, but they don’t.

NAACL 2022long

When strong partial-input baselines reveal artifacts in crowdsourced NLI datasets, the performance of full-input models trained on such datasets is often dismissed as reliance on spurious correlations. We investigate whether state-of-the-art NLI models are capable of overriding default inferences ma…

2022

Recognition of They/Them as Singular Personal Pronouns in Coreference Resolution

NAACL 2022long

As using they/them as personal pronouns becomes increasingly common in English, it is important that coreference resolution systems work as well for individuals who use personal “they” as they do for those who use gendered personal pronouns. We introduce a new benchmark for coreference resolution sy…

2022

Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models

NAACL 2022long

NLP models trained on text have been shown to reproduce human stereotypes, which can magnify harms to marginalized groups when systems are deployed at scale. We adapt the Agency-Belief-Communion (ABC) stereotype model of Koch et al. (2016) from social psychology as a framework for the systematic stu…

2021

Learning to Rationalize for Nonmonotonic Reasoning with Distant Supervision

AAAI 2021technical

The black-box nature of neural models has motivated a line of research that aims to generate natural language rationales to explain why a model made certain predictions. Such rationale generation models, to date, have been trained on dataset-specific crowdsourced rationales, but this approach is cos…

Cited by 39SourcePDFScholar
2021

MedNLI Is Not Immune: Natural Language Inference Artifacts in the Clinical Domain

ACL 2021short

Crowdworker-constructed natural language inference (NLI) datasets have been found to contain statistical artifacts associated with the annotation process that allow hypothesis-only classifiers to achieve better-than-random performance (CITATION). We investigate whether MedNLI, a physician-annotated…