← Search

Rada Mihalcea

71 accepted papers

2026

Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping

AAAI 2026technical

Culture shapes the objects people use and for what purposes, yet mainstream Vision-Language (VL) datasets frequently exhibit cultural biases, disproportionately favoring higher-income, Western contexts. This imbalance reduces model generalizability and perpetuates performance disparities, especially

Cited by 0SourcePDFScholar
2026

Position: Safe Models Do Not Guarantee Safe Societies: The Case for Sociopolitical Risk

ICML 2026spotlight

Sociopolitical AI risks are threats to collective self-determination: a society's capacity to articulate its interests and realize them through institutions. We argue that sociopolitical AI risks emerge when general-purpose AI systems are integrated into society in ways that disproportionately ampli…

Cited by 0SourceScholar
2025

Acoustic Individual Identification of White-Faced Capuchin Monkeys Using Joint Multi-Species Embeddings

ACL 2025short

Acoustic individual identification of wild animals is an essential task for understanding animal vocalizations within their social contexts, and for facilitating conservation and wildlife monitoring efforts. However, most of the work in this space relies on human efforts, as the development of metho…

Cited by 0SourcePDFScholar
2025

Benchmarking and Improving LLM Robustness for Personalized Generation

EMNLP 2025

Recent years have witnessed a growing interest in personalizing the responses of large language models (LLMs). While existing evaluations primarily focus on whether a response aligns with a user’s preferences, we argue that factuality is an equally important yet often overlooked dimension. In the co

Cited by 0SourcePDFScholar
2025

Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias

NAACL 2025findings

Generated texts from large language models (LLMs) have been shown to exhibit a variety of harmful, human-like biases against various demographics. These findings motivate research efforts aiming to understand and measure such effects. This paper introduces a causal formulation for bias measurement i…

2025

Chumor 2.0: Towards Better Benchmarking Chinese Humor Understanding from (Ruo Zhi Ba)

ACL 2025finding

Existing humor datasets and evaluations predominantly focus on English, leaving limited resources for culturally nuanced humor in non-English languages like Chinese. To address this gap, we construct **Chumor**, the first and the largest Chinese humor explanation dataset. **Chumor** is sourced from…

2025

CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation

ACL 2025finding

In clinical operations, teamwork can be the crucial factor that determines the final outcome. Prior studies have shown that sufficient collaboration is the key factor that determines the outcome of an operation. To understand how the team practices teamwork during the operation, we collected **Clini…

2025

Eeyore: Realistic Depression Simulation via Expert-in-the-Loop Supervised and Preference Optimization

ACL 2025finding

Large Language Models (LLMs) have been previously explored for mental healthcare training and therapy client simulation, but they still fall short in authentically capturing diverse client traits and psychological conditions. We introduce Eeyore , an 8B model optimized for realistic depression simul…

2025

Evaluating LLMs’ Mathematical and Coding Competency through Ontology-guided Interventions

ACL 2025finding

Recent advancements in Large Language Models (LLMs) have showcased striking results on existing logical reasoning benchmarks, with some models even surpassing human performance. However, the true depth of their competencies and robustness in reasoning tasks remains an open question. To this end, in…

2025

Examining Spanish Counseling with MIDAS: a Motivational Interviewing Dataset in Spanish

NAACL 2025short

Cultural and language factors significantly influence counseling, but Natural Language Processing research has not yet examined whether the findings of conversational analysis for counseling conducted in English apply to other languages. This paper presents a first step towards this direction. We in…

2025

Language Model Alignment in Multilingual Trolley Problems

ICLR 2025spotlight

We evaluate the moral alignment of large language models (LLMs) with human preferences in multilingual trolley problems. Building on the Moral Machine experiment, which captures over 40 million human judgments across 200+ countries, we develop a cross-lingual corpus of moral dilemma vignettes in ove…

Cited by 3SourcePDFScholar
2025

MAiDE-up: Multilingual Deception Detection of AI-generated Hotel Reviews

NAACL 2025findings

Deceptive reviews are becoming increasingly common, especially given the increase in performance and the prevalence of LLMs. While work to date has addressed the development of models to differentiate between truthful and deceptive human reviews, much less is known about the distinction between real…

2025

Mind the (Belief) Gap: Group Identity in the World of LLMs

ACL 2025finding

Social biases and belief-driven behaviors can significantly impact Large Language Models’ (LLMs’) decisions on several tasks. As LLMs are increasingly used in multi-agent systems for societal simulations, their ability to model fundamental group psychological characteristics remains critical yet und…

2025

MoMentS: A Comprehensive Multimodal Benchmark for Theory of Mind

EMNLP 2025

Understanding Theory of Mind is essential for building socially intelligent multimodal agents capable of perceiving and interpreting human behavior. We introduce MoMentS (Multimodal Mental States), a comprehensive benchmark designed to assess the ToM capabilities of multimodal large language models

2025

Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing

ACL 2025finding

The ability of Natural Language Processing (NLP) methods to categorize text into multiple classes has motivated their use in online content moderation tasks, such as hate speech and fake news detection. However, there is limited understanding of how or why these methods make such decisions, or why c…

2025

Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?

EMNLP 2025

The value orientation of Large Language Models (LLMs) has been extensively studied, as it can shape user experiences across demographic groups.However, two key challenges remain: (1) the lack of systematic comparison across value probing strategies, despite the Multiple Choice Question (MCQ) setting

Cited by 0SourcePDFScholar
2025

The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning

NAACL 2025long

Large Multimodal Models (LMMs) exhibit impressive performance across various multimodal tasks. However, their effectiveness in cross-cultural contexts remains limited due to the predominantly Western-centric nature of most data and models. Conversely, multi-agent models have shown significant capabi…

2025

Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models

NAACL 2025long

Recent work has demonstrated that the unequal representation of cultures and socioeconomic groups in training data leads to biased Large Multi-modal (LMM) models. To improve LMM model performance on underrepresented data, we propose and evaluate several prompting strategies using non-English, geogra…

2025

Why AI Is WEIRD and Shouldn't Be This Way: Towards AI for Everyone, with Everyone, by Everyone

AAAI 2025technical

This paper presents a vision for creating AI systems that are inclusive at every stage of development, from data collection to model design and evaluation. We address key limitations in the current AI pipeline and its WEIRD* representation, such as lack of data diversity, biases in model performance…

Cited by 5SourcePDFScholar
2024

A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

ICML 2024oral

While alignment algorithms are commonly used to tune pre-trained language models towards user preferences, we lack explanations for the underlying mechanisms in which models become ``aligned'', thus making it difficult to explain phenomena like jailbreaks. In this work we study a popular algorithm,…

2024

Analyzing Occupational Distribution Representation in Japanese Language Models

COLING 2024main

Recent advances in large language models (LLMs) have enabled users to generate fluent and seemingly convincing text. However, these models have uneven performance in different languages, which is also associated with undesirable societal biases toward marginalized populations. Specifically, there is…

Cited by 1SourcePDFScholar
2024

Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost

COLING 2024main

Current foundation models have shown impressive performance across various tasks. However, several studies have revealed that these models are not effective for everyone due to the imbalanced geographical and economic representation of the data used in the training process. Most of this data comes f…

2024

CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

NeurIPS 2024oral

Visual Question Answering~(VQA) is an important task in multimodal AI, which requires models to understand and reason on knowledge present in visual and textual data. However, most of the current VQA datasets and models are primarily focused on English and a few major world languages, with images th…

Cited by 34SourcePDFScholar
2024

Can Large Language Models Infer Causation from Correlation?

ICLR 2024poster

Causal inference is one of the hallmarks of human intelligence. While the field of CausalNLP has attracted much interest in the recent years, existing causal inference datasets in NLP primarily rely on discovering causality from empirical knowledge (e.g., commonsense knowledge). In this work, we pro…

2024

Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents

NeurIPS 2024poster

As AI systems pervade human life, ensuring that large language models (LLMs) make safe decisions remains a significant challenge. We introduce the Governance of the Commons Simulation (GovSim), a generative simulation platform designed to study strategic interactions and cooperative decision-making…

2024

Do LLMs Think Fast and Slow? A Causal Study on Sentiment Analysis

EMNLP 2024finding

Sentiment analysis (SA) aims to identify the sentiment expressed in a piece of text, often in the form of a review. Assuming a review and the sentiment associated with it, in this paper we formulate SA as a combination of two tasks: (1) a causal discovery task that distinguishes whether a review “pr…

2024

Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation

COLING 2024main

In this paper, we study the problem of multi-reward reinforcement learning to jointly optimize for multiple text qualities for natural language generation. We focus on the task of counselor reflection generation, where we optimize the generators to simultaneously improve the fluency, coherence, and…

2024

EmoBench: Evaluating the Emotional Intelligence of Large Language Models

ACL 2024long

Recent advances in Large Language Models (LLMs) have highlighted the need for robust, comprehensive, and challenging benchmarks. Yet, research on evaluating their Emotional Intelligence (EI) is considerably limited. Existing benchmarks have two major shortcomings: first, they mainly focus on emotion…

2024

Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language Models

COLING 2024main

Recent progress in large language models (LLMs) has enabled the deployment of many generative NLP applications. At the same time, it has also led to a misleading public discourse that “it’s all been solved.” Not surprisingly, this has, in turn, made many NLP researchers – especially those at the beg…

Cited by 9SourcePDFScholar
2024

Implicit Personalization in Language Models: A Systematic Study

EMNLP 2024finding

Implicit Personalization (IP) is a phenomenon of language models inferring a user’s background from the implicit cues in the input prompts and tailoring the response based on this inference. While previous work has touched upon various instances of this problem, there lacks a unified framework to st…

2024

Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs

ACL 2024findings

Tables contrast with unstructured text data by its structure to organize the information.In this paper, we investigate the efficiency of various LLMs in interpreting tabular data through different prompting strategies and data formats. Our analysis extends across six benchmarks for table-related tas…

Cited by 10SourcePDFScholar
2024

The Generation Gap: Exploring Age Bias in the Value Systems of Large Language Models

EMNLP 2024main

We explore the alignment of values in Large Language Models (LLMs) with specific age groups, leveraging data from the World Value Survey across thirteen categories. Through a diverse set of prompts tailored to ensure response robustness, we find a general inclination of LLM values towards younger de…

2024

Towards Algorithmic Fidelity: Mental Health Representation across Demographics in Synthetic vs. Human-generated Data

COLING 2024main

Synthetic data generation has the potential to impact applications and domains with scarce data. However, before such data is used for sensitive tasks such as mental health, we need an understanding of how different demographics are represented in it. In our paper, we analyze the potential of produc…

2024

Towards Dog Bark Decoding: Leveraging Human Speech Processing for Automated Bark Classification

COLING 2024main

Similar to humans, animals make extensive use of verbal and non-verbal forms of communication, including a large range of audio signals. In this paper, we address dog vocalizations and explore the use of self-supervised speech representation models pre-trained on human speech to address dog bark cla…

Cited by 4SourcePDFScholar
2024

Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions

EMNLP 2024finding

As Large Language Models (LLMs) continue to evolve, they are increasingly being employed in numerous studies to simulate societies and execute diverse social tasks. However, LLMs are susceptible to societal biases due to their exposure to human-generated data. Given that LLMs are being used to gain…

2024

Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense

NAACL 2024long

Large language models (LLMs) have demonstrated substantial commonsense understanding through numerous benchmark evaluations. However, their understanding of cultural commonsense remains largely unexamined. In this paper, we conduct a comprehensive examination of the capabilities and limitations of s…

Cited by 33SourcePDFScholar
2023

Bridging the Digital Divide: Performance Variation across Socio-Economic Factors in Vision-Language Models

EMNLP 2023long main

Despite the impressive performance of current AI models reported across various tasks, performance reports often do not include evaluations of how these models perform on the specific groups that will be impacted by these technologies. Among the minority groups under-represented in AI, data from low…

Cited by 0SourcecodeScholar
2023

Evaluating Parameter-Efficient Transfer Learning Approaches on SURE Benchmark for Speech Understanding

ICASSP 2023accepted

Fine-tuning is widely used as the default algorithm for transfer learning from pre-trained models. Parameter inefficiency can however arise when, during transfer learning, all the parameters of a large pre-trained model need to be updated for individual downstream tasks. As the number of parameters…

Cited by 0SourceScholar
2023

Hi-ToM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models

EMNLP 2023long findings

Theory of Mind (ToM) is the ability to reason about one's own and others' mental states. ToM plays a critical role in the development of intelligence, language understanding, and cognitive processes. While previous work has primarily focused on first and second-order ToM, we explore higher-order ToM…

Cited by 0SourcecodeScholar
2023

Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts

EMNLP 2023short findings

Visual question answering (VQA) is the task of answering questions about an image. The task assumes an understanding of both the image and the question to provide a natural language answer. VQA has gained popularity in recent years due to its potential applications in a wide range of fields, includi…

Cited by 0SourcecodeScholar
2023

Task-Adaptive Tokenization: Enhancing Long-Form Text Generation Efficacy in Mental Health and Beyond

EMNLP 2023long main

We propose task-adaptive tokenization\footnote{Our work will be publicly available upon acceptance.} as a way to adapt the generation pipeline to the specifics of a downstream task and enhance long-form generation in mental health. Inspired by insights from cognitive science, our task-adaptive token…

Cited by 0SourcecodeScholar
2023

VERVE: Template-based ReflectiVE Rewriting for MotiVational IntErviewing

EMNLP 2023long findings

Reflective listening is a fundamental skill that counselors must acquire to achieve proficiency in motivational interviewing (MI). It involves responding in a manner that acknowledges and explores the meaning of what the client has expressed in the conversation. In this work, we introduce the task o…

Cited by 0SourceScholar
2023

You Are What You Annotate: Towards Better Models through Annotator Representations

EMNLP 2023long findings

Annotator disagreement is ubiquitous in natural language processing (NLP) tasks. There are multiple reasons for such disagreements, including the subjectivity of the task, difficult cases, unclear guidelines, and so on. Rather than simply aggregating labels to obtain data annotations, we instead try…

Cited by 0SourcecodeScholar
2022

CICERO: A Dataset for Contextualized Commonsense Inference in Dialogues

ACL 2022long

This paper addresses the problem of dialogue reasoning with contextualized commonsense inference. We curate CICERO, a dataset of dyadic conversations with five types of utterance-level reasoning-based inferences: cause, subsequent event, prerequisite, motivation, and emotional reaction. The dataset…

2022

FIBER: Fill-in-the-Blanks as a Challenging Video Understanding Evaluation Framework

ACL 2022long

We propose fill-in-the-blanks as a video understanding evaluation framework and introduce FIBER – a novel dataset consisting of 28,000 videos and descriptions in support of this evaluation framework. The fill-in-the-blanks setting tests a model’s understanding of a video by requiring it to predict a…

2022

In-the-Wild Video Question Answering

COLING 2022main

Existing video understanding datasets mostly focus on human interactions, with little attention being paid to the “in the wild” settings, where the videos are recorded outdoors. We propose WILDQA, a video understanding dataset of videos recorded in outside settings. In addition to video question ans…

Cited by 0SourcePDFScholar
2022

Knowledge Enhanced Reflection Generation for Counseling Dialogues

ACL 2022long

In this paper, we study the effect of commonsense and domain knowledge while generating responses in counseling conversations using retrieval and generative methods for knowledge integration. We propose a pipeline that collects domain knowledge through web mining, and show that retrieval from both d…

2022

Leveraging Similar Users for Personalized Language Modeling with Limited Data

ACL 2022long

Personalized language models are designed and trained to capture language patterns specific to individual users. This makes them more accurate at predicting what a user will write. However, when a new user joins a platform and not enough text is available, it is harder to build effective personalize…

Cited by 35SourcePDFScholar
2022

Logical Fallacy Detection

EMNLP 2022finding

Reasoning is central to human intelligence. However, fallacious arguments are common, and some exacerbate problems such as spreading misinformation about climate change. In this paper, we propose the task of logical fallacy detection, and provide a new dataset (Logic) of logical fallacies generally…

2022

PAIR: Prompt-Aware margIn Ranking for Counselor Reflection Scoring in Motivational Interviewing

EMNLP 2022main

Counselor reflection is a core verbal skill used by mental health counselors to express understanding and affirmation of the client’s experience and concerns. In this paper, we propose a system for the analysis of counselor reflections. Specifically, our system takes as input one dialog turn contain…

Cited by 17SourcePDFScholar
2022

Two is Better than Many? Binary Classification as an Effective Approach to Multi-Choice Question Answering

EMNLP 2022main

We propose a simple refactoring of multi-choice question answering (MCQA) tasks as a series of binary classifications. The MCQA task is generally performed by scoring each (question, answer) pair normalized over all the pairs, and then selecting the answer from the pair that yield the highest score.…

2022

Using Paraphrases to Study Properties of Contextual Embeddings

NAACL 2022long

We use paraphrases as a unique source of data to analyze contextualized embeddings, with a particular focus on BERT. Because paraphrases naturally encode consistent word and phrase semantics, they provide a unique lens for investigating properties of embeddings. Using the Paraphrase Database’s align…

Cited by 6SourcePDFScholar
2022

When to Make Exceptions: Exploring Language Models as Accounts of Human Moral Judgment

NeurIPS 2022accept

AI systems are becoming increasingly intertwined with human life. In order to effectively collaborate with humans and ensure safety, AI systems need to be able to understand, interpret and predict human moral judgments and decisions. Human moral judgments are often guided by rules, but not always. A…

2021

Analyzing the Surprising Variability in Word Embedding Stability Across Languages

EMNLP 2021main

Word embeddings are powerful representations that form the foundation of many natural language processing architectures, both in English and in other languages. To gain further insight into word embeddings, we explore their stability (e.g., overlap between the nearest neighbors of a word in differen…

2021

Hitting your MARQ: Multimodal ARgument Quality Assessment in Long Debate Video

EMNLP 2021main

The combination of gestures, intonations, and textual content plays a key role in argument delivery. However, the current literature mostly considers textual content while assessing the quality of an argument, and it is limited to datasets containing short sequences (18-48 words). In this paper, we…

Cited by 5SourcePDFScholar
2021

Humor Knowledge Enriched Transformer for Understanding Multimodal Humor

AAAI 2021technical

Recognizing humor from a video utterance requires understanding the verbal and non-verbal components as well as incorporating the appropriate context and external knowledge. In this paper, we propose Humor Knowledge enriched Transformer (HKT) that can capture the gist of a multimodal humorous expres…

2021

MUSER: MUltimodal Stress detection using Emotion Recognition as an Auxiliary Task

NAACL 2021long

The capability to automatically detect human stress can benefit artificial intelligent agents involved in affective computing and human-computer interaction. Stress and emotion are both human affective states, and stress has proven to have important implications on the regulation and expression of e…

Cited by 18SourcePDFScholar
2021

Micromodels for Efficient, Explainable, and Reusable Systems: A Case Study on Mental Health

EMNLP 2021finding

Many statistical models have high accuracy on test benchmarks, but are not explainable, struggle in low-resource scenarios, cannot be reused for multiple tasks, and cannot easily integrate domain expertise. These factors limit their use, particularly in settings such as mental health, where it is di…

2021

Mining the Cause of Political Decision-Making from Social Media: A Case Study of COVID-19 Policies across the US States

EMNLP 2021finding

Mining the causes of political decision-making is an active research area in the field of political science. In the past, most studies have focused on long-term policies that are collected over several decades of time, and have primarily relied on surveys as the main source of predictors. However, t…

Cited by 16SourcePDFScholar
2021

STaCK: Sentence Ordering with Temporal Commonsense Knowledge

EMNLP 2021main

Sentence order prediction is the task of finding the correct order of sentences in a randomly ordered document. Correctly ordering the sentences requires an understanding of coherence with respect to the chronological sequence of events described in the text. Document-level contextual understanding…

2021

WhyAct: Identifying Action Reasons in Lifestyle Vlogs

EMNLP 2021main

We aim to automatically identify human action reasons in online videos. We focus on the widespread genre of lifestyle vlogs, in which people perform actions while verbally describing them. We introduce and make publicly available the WhyAct dataset, consisting of 1,077 visual actions manually annota…

2020

Biased TextRank: Unsupervised Graph-Based Content Extraction

COLING 2020main

We introduce Biased TextRank, a graph-based content extraction method inspired by the popular TextRank algorithm that ranks text spans according to their importance for language processing tasks and according to their relevance to an input “focus.” Biased TextRank enables focused content extraction…

Cited by 43SourcePDFScholar
2020

Exploring the Value of Personalized Word Embeddings

COLING 2020main

In this paper, we introduce personalized word embeddings, and examine their value for language modeling. We compare the performance of our proposed prediction model when using personalized versus generic word representations, and study how these representations can be leveraged for improved performa…

Cited by 19SourcePDFScholar
2020

“Judge me by my size (noun), do you?” YodaLib: A Demographic-Aware Humor Generation Framework

COLING 2020main

The subjective nature of humor makes computerized humor generation a challenging task. We propose an automatic humor generation framework for filling the blanks in Mad Libs® stories, while accounting for the demographic backgrounds of the desired audience. We collect a dataset consisting of such sto…

Cited by 5SourcePDFScholar
2019

Muse-ing on the Impact of Utterance Ordering on Crowdsourced Emotion Annotations

ICASSP 2019accepted

Emotion recognition algorithms rely on data annotated with high quality labels. However, emotion expression and perception are inherently subjective. There is generally not a single annotation that can be unambiguously declared "correct." As a result, annotations are colored by the manner in which t…

Cited by 0SourceScholar