← Search

Shreya Havaldar

7 accepted papers

2025

Adaptively profiling models with task elicitation

EMNLP 2025

Language model evaluations often fail to characterize consequential failure modes, forcing experts to inspect outputs and build new benchmarks. We introduce task elicitation, a method that automatically builds new evaluations to profile model behavior. Task elicitation finds hundreds of natural-lang

2025

Entailed Between the Lines: Incorporating Implication into NLI

ACL 2025long

Much of human communication depends on implication, conveying meaning beyond literal words to express a wider range of thoughts, intentions, and feelings. For models to better understand and facilitate human communication, they must be responsive to the text’s implicit meaning. We focus on Natural L…

2025

Probabilistic Soundness Guarantees in LLM Reasoning Chains

EMNLP 2025

In reasoning chains generated by large language models (LLMs), initial errors often propagate and undermine the reliability of the final conclusion. Current LLM-based error detection methods often fail to detect propagated errors because earlier errors can corrupt judgments of downstream reasoning.

2025

Social Norms in Cinema: A Cross-Cultural Analysis of Shame, Pride and Prejudice

NAACL 2025long

Shame and pride are social emotions expressed across cultures to motivate and regulate people’s thoughts, feelings, and behaviors. In this paper, we introduce the first cross-cultural dataset of over 10k shame/pride-related expressions with underlying social expectations from ~5.4K Bollywood and Hol…

Cited by 1SourcePDFScholar
2024

Building Knowledge-Guided Lexica to Model Cultural Variation

NAACL 2024long

Cultural variation exists between nations (e.g., the United States vs. China), but also within regions (e.g., California vs. Texas, Los Angeles vs. San Francisco). Measuring this regional cultural variation can illuminate how and why people think and behave differently. Historically, it has been dif…