← Search

Caleb Ziems

17 accepted papers

2025

Culture Cartography: Mapping the Landscape of Cultural Knowledge

EMNLP 2025

To serve global users safely and productively, LLMs need culture-specific knowledge that might not be learned during pre-training. How do we find knowledge that is (1) salient to in-group users, but (2) unknown to LLMs? The most common solutions are single-initiative: either researchers define chall

2025

EgoNormia: Benchmarking Physical-Social Norm Understanding

ACL 2025finding

Human activity is moderated by norms; however, supervision for normative reasoning is sparse, particularly where norms are physically- or socially-grounded. We thus present EgoNormia \lVert 𝜖 \rVert, comprising 1,853 (200 for EgoNormia-verified) multiple choice questions (MCQs) grounded within ego-c…

2024

CultureBank: An Online Community-Driven Knowledge Base Towards Culturally Aware Language Technologies

EMNLP 2024finding

To enhance language models’ cultural awareness, we design a generalizable pipeline to construct cultural knowledge bases from different online communities on a massive scale. With the pipeline, we construct CultureBank, a knowledge base built upon users’ self-narratives with 12K cultural descriptors…

2024

Measuring and Addressing Indexical Bias in Information Retrieval

ACL 2024findings

Information Retrieval (IR) systems are designed to deliver relevant content, but traditional systems may not optimize rankings for fairness, neutrality, or the balance of ideas. Consequently, IR can often introduce indexical biases, or biases in the positional order of documents. Although indexical…

2024

Silent Signals, Loud Impact: LLMs for Word-Sense Disambiguation of Coded Dog Whistles

ACL 2024long

A dog whistle is a form of coded communication that carries a secondary meaning to specific audiences and is often weaponized for racial and socioeconomic discrimination. Dog whistling historically originated from United States politics, but in recent years has taken root in social media as a means…

Cited by 1SourcePDFScholar
2024

Social Intelligence Data Infrastructure: Structuring the Present and Navigating the Future

ACL 2024findings

As Natural Language Processing (NLP) systems become increasingly integrated into human social life, these technologies will need to increasingly rely on social intelligence. Although there are many valuable datasets that benchmark isolated dimensions of social intelligence, there does not yet exist…

Cited by 6SourcePDFScholar
2023

CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation

EMNLP 2023long main

Annotated data plays a critical role in Natural Language Processing (NLP) in training models and evaluating their performance. Given recent developments in Large Language Models (LLMs), models such as ChatGPT demonstrate zero-shot capability on many text-annotation tasks, comparable with or even exc…

Cited by 0SourcecodeScholar
2023

Modeling Cross-Cultural Pragmatic Inference with Codenames Duet

ACL 2023findings

Pragmatic reference enables efficient interpersonal communication. Prior work uses simple reference games to test models of pragmatic reasoning, often with unidentified speakers and listeners. In practice, however, speakers’ sociocultural background shapes their pragmatic assumptions. For example, r…

2023

Multi-VALUE: A Framework for Cross-Dialectal English NLP

ACL 2023long

Dialect differences caused by regional, social, and economic factors cause performance discrepancies for many groups of language technology users. Inclusive and equitable language technology must critically be dialect invariant, meaning that performance remains constant over dialectal shifts. Curren…

Cited by 46SourcePDFScholar
2023

NormBank: A Knowledge Bank of Situational Social Norms

ACL 2023long

We present NormBank, a knowledge bank of 155k situational norms. This resource is designed to ground flexible normative reasoning for interactive, assistive, and collaborative AI systems. Unlike prior commonsense resources, NormBank grounds each inference within a multivalent sociocultural frame, wh…

2022

The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems

ACL 2022long

Conversational agents have come increasingly closer to human competence in open-domain dialogue settings; however, such models can reflect insensitive, hurtful, or entirely incoherent viewpoints that erode a user’s trust in the moral integrity of the system. Moral deviations are difficult to mitigat…

2022

VALUE: Understanding Dialect Disparity in NLU

ACL 2022long

English Natural Language Understanding (NLU) systems have achieved great performances and even outperformed humans on benchmarks like GLUE and SuperGLUE. However, these benchmarks contain only textbook Standard American English (SAE). Other dialects have been largely overlooked in the NLP community.…

2021

Latent Hatred: A Benchmark for Understanding Implicit Hate Speech

EMNLP 2021main

Hate speech has grown significantly on social media, causing serious consequences for victims of all demographics. Despite much attention being paid to characterize and detect discriminatory speech, most work has focused on explicit or overt hate speech, failing to address a more pervasive form base…