← Search

Dirk Hovy

31 accepted papers

2026

SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors

ICLR 2026poster

Large language model (LLM) simulations of human behavior have the potential to revolutionize the social and behavioral sciences, if and only if they faithfully reflect real human behaviors. Current evaluations are fragmented, based on bespoke tasks and metrics, creating a patchwork of incomparable r…

Cited by 0SourceScholar
2025

Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals’ Subjective Text Perceptions

ACL 2025long

People naturally vary in their annotations for subjective questions and some of this variation is thought to be due to the person’s sociodemographic characteristics. LLMs have also been used to label data, but recent work has shown that models perform poorly when prompted with sociodemographic attri…

Cited by 0SourcePDFScholar
2025

Biased Tales: Cultural and Topic Bias in Generating Children’s Stories

EMNLP 2025

Stories play a pivotal role in human communication, shaping beliefs and morals, particularly in children. As parents increasingly rely on large language models (LLMs) to craft bedtime stories, the presence of cultural and gender stereotypes in these narratives raises significant concerns. To address

2025

SafetyPrompts: A Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety

AAAI 2025technical

The last two years have seen a rapid growth in concerns around the safety of large language models (LLMs). Researchers and practitioners have met these concerns by creating an abundance of datasets for evaluating and improving LLM safety. However, much of this work has happened in parallel, and with…

2025

The AI Gap: How Socioeconomic Status Affects Language Technology Interactions

ACL 2025long

Socioeconomic status (SES) fundamentally influences how people interact with each other and, more recently, with digital technologies like large language models (LLMs). While previous research has highlighted the interaction between SES and language technology, it was limited by reliance on proxy me…

Cited by 0SourcePDFScholar
2024

Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution

ACL 2024long

Large language models (LLMs) reflect societal norms and biases, especially about gender. While societal biases and stereotypes have been extensively researched in various NLP applications, there is a surprising gap for emotion analysis. However, emotion and gender are closely linked in societal disc…

2024

Classist Tools: Social Class Correlates with Performance in NLP

ACL 2024long

The field of sociolinguistics has studied factors affecting language use for the last century. Labov (1964) and Bernstein (1960) showed that socioeconomic class strongly influences our accents, syntax and lexicon. However, despite growing concerns surrounding fairness and bias in Natural Language Pr…

2024

DADIT: A Dataset for Demographic Classification of Italian Twitter Users and a Comparison of Prediction Methods

COLING 2024main

Social scientists increasingly use demographically stratified social media data to study the attitudes, beliefs, and behavior of the general public. To facilitate such analyses, we construct, validate, and release publicly the representative DADIT dataset of 30M tweets of 20k Italian Twitter users,…

2024

Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models

EMNLP 2024finding

Emotions play important epistemological and cognitive roles in our lives, revealing our values and guiding our actions. Previous work has shown that LLMs display biases in emotion attribution along gender lines. However, unlike gender, which says little about our values, religion, as a socio-cultura…

Cited by 6SourcePDFScholar
2024

Emotion Analysis in NLP: Trends, Gaps and Roadmap for Future Directions

COLING 2024main

Emotions are a central aspect of communication. Consequently, emotion analysis (EA) is a rapidly growing field in natural language processing (NLP). However, there is no consensus on scope, direction, or methods. In this paper, we conduct a thorough review of 154 relevant NLP publications from the l…

2024

Impoverished Language Technology: The Lack of (Social) Class in NLP

COLING 2024main

Since Labov’s foundational 1964 work on the social stratification of language, linguistics has dedicated concerted efforts towards understanding the relationships between socio-demographic factors and language production and perception. Despite the large body of evidence identifying significant rela…

Cited by 2SourcePDFScholar
2024

Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models

ACL 2024long

Much recent work seeks to evaluate values and opinions in large language models (LLMs) using multiple-choice surveys and questionnaires. Most of this work is motivated by concerns around real-world LLM applications. For example, politically-biased LLMs may subtly influence society when they are used…

2024

Twists, Humps, and Pebbles: Multilingual Speech Recognition Models Exhibit Gender Performance Gaps

EMNLP 2024main

Current automatic speech recognition (ASR) models are designed to be used across many languages and tasks without substantial changes. However, this broad language coverage hides performance gaps within languages, for example, across genders. Our study systematically evaluates the performance of two…

2024

XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

NAACL 2024long

Without proper safeguards, large language models will readily follow malicious instructions and generate toxic content. This risk motivates safety efforts such as red-teaming and large-scale feedback learning, which aim to make models both helpful and harmless. However, there is a tension between th…

2024

“My Answer is C”: First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models

ACL 2024findings

The open-ended nature of language generation makes the evaluation of autoregressive large language models (LLMs) challenging. One common evaluation approach uses multiple-choice questions to limit the response space. The model is then evaluated by ranking the candidate answers by the log probability…

2023

The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics

ACL 2023short

Many NLP tasks exhibit human label variation, where different annotators give different labels to the same texts. This variation is known to depend, at least in part, on the sociodemographics of annotators. Recent research aims to model individual annotator behaviour rather than predicting aggregate…

Cited by 24SourcePDFScholar
2023

The State of Profanity Obfuscation in Natural Language Processing Scientific Publications

ACL 2023findings

Work on hate speech has made considering rude and harmful examples in scientific publications inevitable. This situation raises various problems, such as whether or not to obscure profanities. While science must accurately disclose what it does, the unwarranted spread of hate speech can harm readers…

Cited by 16SourcePDFScholar
2023

What about “em”? How Commercial Machine Translation Fails to Handle (Neo-)Pronouns

ACL 2023long

As 3rd-person pronoun usage shifts to include novel forms, e.g., neopronouns, we need more research on identity-inclusive NLP. Exclusion is particularly harmful in one of the most popular NLP applications, machine translation (MT). Wrong pronoun translations can discriminate against marginalized gro…

Cited by 23SourcePDFScholar
2022

Bridging Fairness and Environmental Sustainability in Natural Language Processing

EMNLP 2022main

Fairness and environmental impact are important research directions for the sustainable development of artificial intelligence. However, while each topic is an active research area in natural language processing (NLP), there is a surprising lack of research on the interplay between the two fields. T…

Cited by 17SourcePDFScholar
2022

Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages

EMNLP 2022main

Hate speech is a global phenomenon, but most hate speech datasets so far focus on English-language content. This hinders the development of more effective hate speech detection models in hundreds of languages spoken by billions across the world. More data is needed, but annotating hateful content is…

2022

Entropy-based Attention Regularization Frees Unintended Bias Mitigation from Lists

ACL 2022findings

Natural Language Processing (NLP) models risk overfitting to specific terms in the training data, thereby reducing their performance, fairness, and generalizability. E.g., neural hate speech detection models are strongly influenced by identity terms like gay, or women, resulting in false positives,…

2022

SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems

ACL 2022long

The social impact of natural language processing and its applications has received increasing attention. In this position paper, we focus on the problem of safety for end-to-end conversational AI. We survey the problem landscape therein, introducing a taxonomy of three observed phenomena: the Instig…

2022

SocioProbe: What, When, and Where Language Models Learn about Sociodemographics

EMNLP 2022main

Pre-trained language models (PLMs) have outperformed other NLP models on a wide range of tasks. Opting for a more thorough understanding of their capabilities and inner workings, researchers have established the extend to which they capture lower-level knowledge like grammaticality, and mid-level se…

2022

Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks

NAACL 2022long

Labelled data is the foundation of most natural language processing tasks. However, labelling data is difficult and there often are diverse valid beliefs about what the correct data labels should be. So far, dataset creators have acknowledged annotator subjectivity, but rarely actively managed it in…

2022

Welcome to the Modern World of Pronouns: Identity-Inclusive Natural Language Processing beyond Gender

COLING 2022main

The world of pronouns is changing – from a closed word class with few members to an open set of terms to reflect identities. However, Natural Language Processing (NLP) barely reflects this linguistic shift, resulting in the possible exclusion of non-binary users, even though recent work outlined the…

2022

“It’s Not Just Hate”: A Multi-Dimensional Perspective on Detecting Harmful Speech Online

EMNLP 2022main

Well-annotated data is a prerequisite for good Natural Language Processing models. Too often, though, annotation decisions are governed by optimizing time or annotator agreement. We make a case for nuanced efforts in an interdisciplinary setting for annotating offensive online speech. Detecting offe…

Cited by 21SourcePDFScholar
2021

Beyond Black & White: Leveraging Annotator Disagreement via Soft-Label Multi-Task Learning

NAACL 2021long

Supervised learning assumes that a ground truth label exists. However, the reliability of this ground truth depends on human annotators, who often disagree. Prior work has shown that this disagreement can be helpful in training models. We propose a novel method to incorporate this disagreement as in…

Cited by 132SourcePDFScholar
2021

HONEST: Measuring Hurtful Sentence Completion in Language Models

NAACL 2021long

Language models have revolutionized the field of NLP. However, language models capture and proliferate hurtful stereotypes, especially in text generation. Our results show that 4.3% of the time, language models complete a sentence with a hurtful word. These cases are not random, but follow language…

2021

Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence

ACL 2021short

Topic models extract groups of words from documents, whose interpretation as a topic hopefully allows for a better understanding of the data. However, the resulting word groups are often not coherent, making them harder to interpret. Recently, neural topic models have shown improvements in overall c…