← Search

Malihe Alikhani

34 accepted papers

2025

Accounting for Sycophancy in Language Model Uncertainty Estimation

NAACL 2025findings

Effective human-machine collaboration requires machine learning models to externalize uncertainty, so users can reflect and intervene when necessary. For language models, these representations of uncertainty may be impacted by sycophancy bias: proclivity to agree with users, even if they are wrong.…

2025

Adaptive Platt Scaling with Causal Interpretations for Self-Reflective Language Model Uncertainty Estimates

EMNLP 2025

As large language models (LLMs) are consumed by more users and deployed in increasingly autonomous capacities, their ability to self-monitor and ask for human intervention is of vital importance. Underlying this capability are fundamental skills like self-reflection and expression of uncertainty. In

2025

An Active Learning Framework for Inclusive Generation by Large Language Models

COLING 2025main

Ensuring that Large Language Models (LLMs) generate text representative of diverse sub-populations is essential, particularly when key concepts related to under-represented groups are scarce in the training data. We address this challenge with a novel clustering-based active learning framework, enha…

2025

Contextual ASR Error Handling with LLMs Augmentation for Goal-Oriented Conversational AI

COLING 2025industry

General-purpose automatic speech recognition (ASR) systems do not always perform well in goal-oriented dialogue. Existing ASR correction methods rely on prior user data or named entities. We extend correction to tasks that have no prior user data and exhibit linguistic flexibility such as lexical an…

2025

Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others from Conversational Cues

ACL 2025long

Typically, when evaluating Theory of Mind, we consider the beliefs of others to be binary: held or not held. But what if someone is unsure about their own beliefs? How can we quantify this uncertainty? We propose a new suite of tasks, challenging language models (LMs) to model the uncertainty of par…

2025

How to Align Multiple Signed Language Corpora for Better Sign-to-Sign Translations?

NAACL 2025long

There are more than 300 documented signed languages worldwide, which are indispensable avenues for computational linguists to study cross-cultural and cross-linguistic factors that affect automatic sign understanding and generation. Yet, these are studied under critically low-resource settings, espe…

2025

Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation

EMNLP 2025

Establishing shared goals is a fundamental step in human-AI communication. However, ambiguities can lead to outputs that seem correct but fail to reflect the speaker’s intent. In this paper, we explore this issue with a focus on the data visualization domain, where ambiguities in natural language im

Cited by 0SourcePDFScholar
2025

Measuring Bias and Agreement in Large Language Model Presupposition Judgments

ACL 2025finding

Identifying linguistic bias in text demands the identification not only of explicitly asserted content but also of implicit content including presuppositions. Large language models (LLMs) offer a promising automated approach to detecting presuppositions, yet the extent to which their judgments align…

Cited by 0SourcePDFScholar
2025

SignAlignLM: Integrating Multimodal Sign Language Processing into Large Language Models

ACL 2025finding

Deaf and Hard-of-Hearing (DHH) users increasingly utilize Large Language Models (LLMs), yet face significant challenges due to these models’ limited understanding of sign language grammar, multimodal sign inputs, and Deaf cultural contexts. Further, current approaches that try to address these limit…

2024

Combining Discourse Coherence with Large Language Models for More Inclusive, Equitable, and Robust Task-Oriented Dialogue

COLING 2024main

Large language models (LLMs) are capable of generating well-formed responses, but using LLMs to generate responses on the fly is not yet feasible for many task-oriented systems. Modular architectures are often still required for safety and privacy guarantees on the output. We hypothesize that an off…

Cited by 2SourcePDFScholar
2024

Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models

ACL 2024findings

Effective interlocutors account for the uncertain goals, beliefs, and emotions of others. But even the best human conversationalist cannot perfectly anticipate the trajectory of a dialogue. How well can language models represent inherent uncertainty in conversations? We propose FortUne Dial, an expa…

2024

Generating Signed Language Instructions in Large-Scale Dialogue Systems

NAACL 2024industry

We introduce a goal-oriented conversational AI system enhanced with American Sign Language (ASL) instructions, presenting the first implementation of such a system on a worldwide multimodal conversational AI platform. Accessible through a touch-based interface, our system receives input from users a…

2024

Seeing Eye-to-Eye: Cross-Modal Coherence Relations Inform Eye-gaze Patterns During Comprehension & Production

COLING 2024main

Context influences how we engage with multimodal documents. Describing and processing the content of images is highly correlated with the goals of the discourse. It is known that these underlying cognitive processes can be tapped into by looking at eye movements, but the connection between discourse…

2024

Studying and Mitigating Biases in Sign Language Understanding Models

EMNLP 2024main

Ensuring that the benefits of sign language technologies are distributed equitably among all community members is crucial. Thus, it is important to address potential biases and inequities that may arise from the design or use of these resources. Crowd-sourced sign language datasets, such as the ASL…

Cited by 0SourcePDFScholar
2023

Learning to Generate Equitable Text in Dialogue from Biased Training Data

ACL 2023long

The ingrained principles of fairness in a dialogue system’s decision-making process and generated responses are crucial for user engagement, satisfaction, and task achievement. Absence of equitable and inclusive principles can hinder the formation of common ground, which in turn negatively impacts t…

2023

MedNgage: A Dataset for Understanding Engagement in Patient-Nurse Conversations

ACL 2023findings

Patients who effectively manage their symptoms often demonstrate higher levels of engagement in conversations and interventions with healthcare practitioners. This engagement is multifaceted, encompassing cognitive and social dimensions. Consequently, it is crucial for AI systems to understand the e…

2023

SODA: Million-scale Dialogue Distillation with Social Commonsense Contextualization

EMNLP 2023long main

Data scarcity has been a long standing issue in the field of open-domain social dialogue. To quench this thirst, we present SODA: the first publicly available, million-scale high-quality social dialogue dataset. By contextualizing social commonsense knowledge from a knowledge graph, we are able to d…

Cited by 0SourcecodeScholar
2022

APPDIA: A Discourse-aware Transformer-based Style Transfer Model for Offensive Social Media Conversations

COLING 2022main

Using style-transfer models to reduce offensiveness of social media comments can help foster a more inclusive environment. However, there are no sizable datasets that contain offensive texts and their inoffensive counterparts, and fine-tuning pretrained models with limited labeled data can lead to t…

2022

Cross-Modal Coherence for Text-to-Image Retrieval

AAAI 2022technical

Common image-text joint understanding techniques presume that images and the associated text can universally be characterized by a single implicit model. However, co-occurring images and text can be related in qualitatively different ways, and explicitly modeling it could improve the performance of…

2022

Modeling Intensification for Sign Language Generation: A Computational Approach

ACL 2022findings

End-to-end sign language generation models do not accurately represent the prosody in sign language. A lack of temporal and spatial variations leads to poor-quality generated presentations that confuse human interpreters. In this paper, we aim to improve the prosody in generated sign languages by mo…

2022

NAREOR: The Narrative Reordering Problem

AAAI 2022technical

Many implicit inferences exist in text depending on how it is structured that can critically impact the text's interpretation and meaning. One such structural aspect present in text with chronology is the order of its presentation. For narratives or stories, this is known as the narrative order. Reo…

2022

PAC-Bayesian domain adaptation bounds for multiclass learners

UAI 2022poster

Multiclass neural networks are a common tool in modern unsupervised domain adaptation, yet an appropriate theoretical description for their non-uniform sample complexity is lacking in the adaptation literature. To fill this gap, we propose the first PAC-Bayesian adaptation bounds for multiclass lear…

2022

PINEAPPLE: Personifying INanimate Entities by Acquiring Parallel Personification Data for Learning Enhanced Generation

COLING 2022main

A personification is a figure of speech that endows inanimate entities with properties and actions typically seen as requiring animacy. In this paper, we explore the task of personification generation. To this end, we propose PINEAPPLE: Personifying INanimate Entities by Acquiring Parallel Personifi…

2022

Political Ideology and Polarization: A Multi-dimensional Approach

NAACL 2022long

Analyzing ideology and polarization is of critical importance in advancing our grasp of modern politics. Recent research has made great strides towards understanding the ideological bias (i.e., stance) of news media along the left-right spectrum. In this work, we instead take a novel and more nuance…

2022

Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation Models

AAAI 2022technical

We investigate the use of multimodal information contained in images as an effective method for enhancing the commonsense of Transformer models for text generation. We perform experiments using BART and T5 on concept-to-text generation, specifically the task of generative commonsense reasoning, or C…

2022

The Change that Matters in Discourse Parsing: Estimating the Impact of Domain Shift on Parser Error

ACL 2022findings

Discourse analysis allows us to attain inferences of a text document that extend beyond the sentence-level. The current performance of discourse models is very low on texts outside of the training distribution’s coverage, diminishing the practical utility of existing models. There is need for a meas…

2022

The Role of Context and Uncertainty in Shallow Discourse Parsing

COLING 2022main

Discourse parsing has proven to be useful for a number of NLP tasks that require complex reasoning. However, over a decade since the advent of the Penn Discourse Treebank, predicting implicit discourse relations in text remains challenging. There are several possible reasons for this, and we hypothe…

Cited by 0SourcePDFScholar
2021

COSMic: A Coherence-Aware Generation Metric for Image Descriptions

EMNLP 2021finding

Developers of text generation models rely on automated evaluation metrics as a stand-in for slow and expensive manual evaluations. However, image captioning metrics have struggled to give accurate learned estimates of the semantic and pragmatic success of output text. We address this weakness by int…

2021

Including Signed Languages in Natural Language Processing

ACL 2021long

Signed languages are the primary means of communication for many deaf and hard of hearing individuals. Since signed languages exhibit all the fundamental linguistic properties of natural language, we believe that tools and theories of Natural Language Processing (NLP) are crucial towards its modelin…

Cited by 130SourcePDFScholar
2020

Aspectuality Across Genre: A Distributional Semantics Approach

COLING 2020main

The interpretation of the lexical aspect of verbs in English plays a crucial role in tasks such as recognizing textual entailment and learning discourse-level inferences. We show that two elementary dimensions of aspectual class, states vs. events, and telic vs. atelic events, can be modelled effect…

Cited by 12SourcePDFScholar
2020

Combining Cognitive Modeling and Reinforcement Learning for Clarification in Dialogue

COLING 2020main

In many domains, dialogue systems need to work collaboratively with users to successfully reconstruct the meaning the user had in mind. In this paper, we show how cognitive models of users’ communicative strategies can be leveraged in a reinforcement learning approach to dialogue planning to enable…

Cited by 13SourcePDFScholar