← Search

Kristina Lerman

14 accepted papers

2025

Aggregation Artifacts in Subjective Tasks Collapse Large Language Models’ Posteriors

NAACL 2025long

In-context Learning (ICL) has become the primary method for performing natural language tasks with Large Language Models (LLMs). The knowledge acquired during pre-training is crucial for this few-shot capability, providing the model with task priors. However, recent studies have shown that ICL predo…

2025

Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts

EMNLP 2025

Modeling complex subjective tasks in Natural Language Processing, such as recognizing emotion and morality, is considerably challenging due to significant variation in human annotations. This variation often reflects reasonable differences in semantic interpretations rather than mere noise, necessit

2025

Improving and Assessing the Fidelity of Large Language Models Alignment to Online Communities

NAACL 2025long

Large language models (LLMs) have shown promise in representing individuals and communities, offering new ways to study complex social dynamics. However, effectively aligning LLMs with specific human groups and systematically assessing the fidelity of the alignment remains a challenge. This paper pr…

2025

Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks

ICASSP 2025accepted

In-Context Learning (ICL) in Large Language Models (LLM) has emerged as the dominant technique for performing natural language tasks, as it does not require updating the model parameters with gradient-based methods. ICL promises to "adapt" the LLM to perform the present task at a competitive or stat…

Cited by 0SourceScholar
2025

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models

EMNLP 2025

Steerability, or the ability of large language models (LLMs) to adapt outputs to align with diverse community-specific norms, perspectives, and communication styles, is critical for real-world applications but remains under-evaluated. We introduce STEER-BENCH, a benchmark for assessing population-sp

Cited by 0SourcePDFScholar
2024

Capturing Perspectives of Crowdsourced Annotators in Subjective Learning Tasks

NAACL 2024long

Supervised classification heavily depends on datasets annotated by humans. However, in subjective tasks such as toxicity classification, these annotations often exhibit low agreement among raters. Annotations have commonly been aggregated by employing methods like majority voting to determine a sing…

2024

Community-Cross-Instruct: Unsupervised Instruction Generation for Aligning Large Language Models to Online Communities

EMNLP 2024main

Social scientists use surveys to probe the opinions and beliefs of populations, but these methods are slow, costly, and prone to biases. Recent advances in large language models (LLMs) enable the creating of computational representations or “digital twins” of populations that generate human-like res…

2024

How Susceptible are Large Language Models to Ideological Manipulation?

EMNLP 2024main

Large Language Models (LLMs) possess the potential to exert substantial influence on public perceptions and interactions with information. This raises concerns about the societal impact that could arise if the ideologies within these models can be easily manipulated. In this work, we investigate how…

2024

Whose Emotions and Moral Sentiments do Language Models Reflect?

ACL 2024findings

Language models (LMs) are known to represent the perspectives of some social groups better than others, which may impact their performance, especially on subjective tasks such as content moderation and hate speech detection. To explore how LMs represent different perspectives, existing research focu…

Cited by 15SourcePDFScholar
2023

Leveraging Label Correlations in a Multi-Label Setting: a Case Study in Emotion

ICASSP 2023accepted

Detecting emotions expressed in text has become critical to a range of fields. In this work, we investigate ways to exploit label correlations in multi-label emotion recognition models to improve emotion detection. First, we develop two modeling approaches to the problem in order to capture word ass…

Cited by 0SourceScholar
2023

Using Emotion Embeddings to Transfer Knowledge between Emotions, Languages, and Annotation Formats

ICASSP 2023accepted

The need for emotional inference from text continues to diversify as more and more disciplines integrate emotions into their theories and applications. These needs include inferring different emotion types, handling multiple languages, and different annotation formats. A shared model between differe…

Cited by 0SourceScholar
2021

Detecting Polarized Topics Using Partisanship-aware Contextualized Topic Embeddings

EMNLP 2021finding

Growing polarization of the news media has been blamed for fanning disagreement, controversy and even violence. Early identification of polarized topics is thus an urgent matter that can help mitigate conflict. However, accurate measurement of topic-wise polarization is still an open research challe…

2021

Speaker Turn Modeling for Dialogue Act Classification

EMNLP 2021finding

Dialogue Act (DA) classification is the task of classifying utterances with respect to the function they serve in a dialogue. Existing approaches to DA classification model utterances without incorporating the turn changes among speakers throughout the dialogue, therefore treating it no different th…

2019

MixHop: Higher-Order Graph Convolutional Architectures via Sparsified Neighborhood Mixing

ICML 2019oral

Existing popular methods for semi-supervised learning with Graph Neural Networks (such as the Graph Convolutional Network) provably cannot learn a general class of neighborhood mixing relationships. To address this weakness, we propose a new model, MixHop, that can learn these relationships, includi…