← Search

David Jurgens

41 accepted papers

2026

SPRIG: Improving Large Language Model Performance by System Prompt Optimization

ICLR 2026poster

Large Language Models (LLMs) have shown impressive capabilities in many scenarios, but their performance depends, in part, on the choice of prompt. Past research has focused on optimizing prompts specific to a task. However, much less attention has been given to optimizing the general instructions i…

Cited by 0SourcecodeScholar
2025

Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs

EMNLP 2025

Personalized Large Language Models (LLMs) are increasingly used in diverse applications, where they are assigned a specific persona—such as a happy high school teacher—to guide their responses. While prior research has examined how well LLMs adhere to predefined personas in writing style, a comprehe

2025

Are Rules Meant to be Broken? Understanding Multilingual Moral Reasoning as a Computational Pipeline with UniMoral

ACL 2025long

Moral reasoning is a complex cognitive process shaped by individual experiences and cultural contexts and presents unique challenges for computational analysis. While natural language processing (NLP) offers promising tools for studying this phenomenon, current research lacks cohesion, employing dis…

2025

Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals’ Subjective Text Perceptions

ACL 2025long

People naturally vary in their annotations for subjective questions and some of this variation is thought to be due to the person’s sociodemographic characteristics. LLMs have also been used to label data, but recent work has shown that models perform poorly when prompted with sociodemographic attri…

Cited by 0SourcePDFScholar
2025

Causally Modeling the Linguistic and Social Factors that Predict Email Response

NAACL 2025long

Email is a vital conduit for human communication across businesses, organizations, and broader societal contexts. In this study, we aim to model the intents, expectations, and responsiveness in email exchanges. To this end, we release SIZZLER, a new dataset containing 1800 emails annotated with nuan…

Cited by 0SourcePDFScholar
2025

Coordinating Chaos: A Structured Review of Linguistic Coordination Methodologies

ACL 2025long

Linguistic coordination—a phenomenon where conversation partners end up having similar patterns of language use—has been established across a variety of contexts and for multiple linguistic features. However, the study of language coordination has been accompanied by a diverse and inconsistently app…

Cited by 0SourcePDFScholar
2025

Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and Domains

EMNLP 2025

Authorship representation (AR) learning, which models an author’s unique writing style, has demonstrated strong performance in authorship attribution tasks. However, prior research has primarily focused on monolingual settings—mostly in English—leaving the potential benefits of multilingual AR model

2025

Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus

ACL 2025long

Podcasts provide highly diverse content to a massive listener base through a unique on-demand modality. However, limited data has prevented large-scale computational analysis of the podcast ecosystem. To fill this gap, we introduce a massive dataset of over 1.1M podcast transcripts that is largely c…

2025

Position: Towards Bidirectional Human-AI Alignment

NeurIPS 2025poster

Recent advances in general-purpose AI underscore the urgent need to align AI systems with human goals and values. Yet, the lack of a clear, shared understanding of what constitutes "alignment" limits meaningful progress and cross-disciplinary collaboration. In this position paper, we argue that the…

Cited by 0SourceScholar
2025

Sociodemographic Prompting is Not Yet an Effective Approach for Simulating Subjective Judgments with LLMs

NAACL 2025short

Human judgments are inherently subjective and are actively affected by personal traits such as gender and ethnicity. While Large LanguageModels (LLMs) are widely used to simulate human responses across diverse contexts, their ability to account for demographic differencesin subjective tasks remains…

2025

Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation Framework

EMNLP 2025

Large language models (LLMs) are increasingly deployed in domains requiring moral understanding, yet their reasoning often remains shallow, and misaligned with human reasoning. Unlike humans, whose moral reasoning integrates contextual trade-offs, value systems, and ethical theories, LLMs often rely

Cited by 0SourcePDFScholar
2025

The Million Authors Corpus: A Cross-Lingual and Cross-Domain Wikipedia Dataset for Authorship Verification

ACL 2025finding

Authorship verification (AV) is a crucial task for applications like identity verification, plagiarism detection, and AI-generated text identification. However, datasets for training and evaluating AV models are primarily in English and primarily in a single domain. This precludes analysis of AV tec…

Cited by 0SourcePDFScholar
2025

The Noisy Path from Source to Citation: Measuring How Scholars Engage with Past Research

ACL 2025long

Academic citations are widely used for evaluating research and tracing knowledge flows. Such uses typically rely on raw citation counts and neglect variability in citation types. In particular, citations can vary in their fidelity as original knowledge from cited studies may be paraphrased, summariz…

2025

The Practical Impacts of Theoretical Constructs on Empathy Modeling

EMNLP 2025

Conceptual operationalizations of empathy in NLP are varied, with some having specific behaviors and properties, while others are more abstract. How these variations relate to one another and capture properties of empathy observable in text remains unclear. To provide insight into this, we analyze t

Cited by 0SourcePDFScholar
2025

Unstructured Evidence Attribution for Long Context Query Focused Summarization

EMNLP 2025

Large language models (LLMs) are capable of generating coherent summaries from very long contexts given a user query, and extracting and citing evidence spans helps improve the trustworthiness of these summaries. Whereas previous work has focused on evidence citation with fixed levels of granularity

2024

Can Large Language Model Agents Simulate Human Trust Behavior?

NeurIPS 2024poster

Large Language Model (LLM) agents have been increasingly adopted as simulation tools to model humans in social science and role-playing applications. However, one fundamental question remains: can LLM agents really simulate human behavior? In this paper, we focus on one critical and elemental behavi…

2024

Finding Educationally Supportive Contexts for Vocabulary Learning with Attention-Based Models

COLING 2024main

When learning new vocabulary, both humans and machines acquire critical information about the meaning of an unfamiliar word through contextual information in a sentence or passage. However, not all contexts are equally helpful for learning an unfamiliar ‘target’ word. Some contexts provide a rich se…

Cited by 0SourcePDFScholar
2024

Tab2Text - A framework for deep learning with tabular data

EMNLP 2024finding

Tabular data, from public opinion surveys to records of interactions with social services, is foundational to the social sciences. One application of such data is to fit supervised learning models in order to predict consequential outcomes, for example: whether a family is likely to be evicted, whet…

2024

The Language of Trauma: Modeling Traumatic Event Descriptions Across Domains with Explainable AI

EMNLP 2024finding

Psychological trauma can manifest following various distressing events and is captured in diverse online contexts. However, studies traditionally focus on a single aspect of trauma, often neglecting the transferability of findings across different scenarios. We address this gap by training various l…

2024

ValueScope: Unveiling Implicit Norms and Values via Return Potential Model of Social Interactions

EMNLP 2024finding

This study introduces ValueScope, a framework leveraging language models to quantify social norms and values within online communities, grounded in social science perspectives on normative structures. We employ ValueScope to dissect and analyze linguistic and stylistic expressions across 13 Reddit c…

2024

When ”A Helpful Assistant” Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models

EMNLP 2024finding

Prompting serves as the major way humans interact with Large Language Models (LLM). Commercial AI systems commonly define the role of the LLM in system prompts. For example, ChatGPT uses ”You are a helpful assistant” as part of its default system prompt. Despite current practices of adding personas…

2024

You don’t need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments

NAACL 2024long

The versatility of Large Language Models (LLMs) on natural language understanding tasks has made them popular for research in social sciences. To properly understand the properties and innate personas of LLMs, researchers have performed studies that involve using prompts in the form of questions tha…

2023

Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark

EMNLP 2023long main

Large language models (LLMs) have been shown to perform well at a variety of syntactic, discourse, and reasoning tasks. While LLMs are increasingly deployed in many forms including conversational agents that interact with humans, we lack a grounded benchmark to measure how well LLMs understand socia…

Cited by 0SourcecodeScholar
2023

When it Rains, it Pours: Modeling Media Storms and the News Ecosystem

EMNLP 2023long findings

Most events in the world receive at most brief coverage by the news media. Occasionally, however, an event will trigger a media storm, with voluminous and widespread coverage lasting for weeks instead of days. In this work, we develop and apply a pairwise article similarity model, allowing us to ide…

Cited by 0SourcecodeScholar
2023

Your spouse needs professional help: Determining the Contextual Appropriateness of Messages through Modeling Social Relationships

ACL 2023long

Understanding interpersonal communication requires, in part, understanding the social context and norms in which a message is said. However, current methods for identifying offensive content in such communication largely operate independent of context, with only a few approaches considering communit…

2022

A Critical Reflection and Forward Perspective on Empathy and Natural Language Processing

EMNLP 2022finding

We review the state of research on empathy in natural language processing and identify the following issues: (1) empathy definitions are absent or abstract, which (2) leads to low construct validity and reproducibility. Moreover, (3) emotional empathy is overemphasized, skewing our focus to a narrow…

Cited by 26SourcePDFScholar
2022

Classification without (Proper) Representation: Political Heterogeneity in Social Media and Its Implications for Classification and Behavioral Analysis

ACL 2022findings

Reddit is home to a broad spectrum of political activity, and users signal their political affiliations in multiple ways—from self-declarations to community participation. Frequently, computational studies have treated political users as a single bloc, both in developing models to infer political le…

Cited by 10SourcePDFScholar
2022

Modeling Information Change in Science Communication with Semantically Matched Paraphrases

EMNLP 2022main

Whether the media faithfully communicate scientific information has long been a core issue to the science community. Automatically identifying paraphrased scientific findings could enable large-scale tracking and analysis of information changes in the science communication process, but this requires…

2022

MultiCite: Modeling realistic citations requires moving beyond the single-sentence single-label setting

NAACL 2022long

Citation context analysis (CCA) is an important task in natural language processing that studies how and why scholars discuss each others’ work. Despite decades of study, computational methods for CCA have largely relied on overly-simplistic assumptions of how authors cite, which ignore several impo…

2021

An animated picture says at least a thousand words: Selecting Gif-based Replies in Multimodal Dialog

EMNLP 2021finding

Online conversations include more than just text. Increasingly, image-based responses such as memes and animated gifs serve as culturally recognized and often humorous responses in conversation. However, while NLP has broadened to multimodal models, conversational dialog systems have largely focused…

2021

Detecting Community Sensitive Norm Violations in Online Conversations

EMNLP 2021finding

Online platforms and communities establish their own norms that govern what behavior is acceptable within the community. Substantial effort in NLP has focused on identifying unacceptable behaviors and, recently, on forecasting them before they occur. However, these efforts have largely focused on to…

2021

Idiosyncratic but not Arbitrary: Learning Idiolects in Online Registers Reveals Distinctive yet Consistent Individual Styles

EMNLP 2021main

An individual’s variation in writing style is often a function of both social and personal attributes. While structured social variation has been extensively studied, e.g., gender based variation, far less is known about how to characterize individual styles due to their idiosyncratic nature. We int…

2021

Using Sociolinguistic Variables to Reveal Changing Attitudes Towards Sexuality and Gender

EMNLP 2021main

Individuals signal aspects of their identity and beliefs through linguistic choices. Studying these choices in aggregate allows us to examine large-scale attitude shifts within a population. Here, we develop computational methods to study word choice within a sociolinguistic lexical variable—alterna…