← Search

Reza Haf

10 accepted papers

2024

An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models

EMNLP 2024main

Large Multimodal Models (LMMs) have achieved strong performance across a range of vision and language tasks. However, their spatial reasoning capabilities are under-investigated. In this paper, we construct a novel VQA dataset, Spatial-MM, to comprehensively study LMMs’ spatial understanding and rea…

2024

Assistive Large Language Model Agents for Socially-Aware Negotiation Dialogues

EMNLP 2024finding

We develop assistive agents based on Large Language Models (LLMs) that aid interlocutors in business negotiations.Specifically, we simulate business negotiations by letting two LLM-based agents engage in role play. A third LLM acts as a remediator agent to rewrite utterances violating norms for impr…

2024

Causal Discovery Inspired Unsupervised Domain Adaptation for Emotion-Cause Pair Extraction

EMNLP 2024finding

This paper tackles the task of emotion-cause pair extraction in the unsupervised domain adaptation setting.The problem is challenging as the distributions of the events causing emotions in target domains are dramatically different than those in source domains, despite the distributions of emotional…

2024

Exploring the Potential of Multimodal LLM with Knowledge-Intensive Multimodal ASR

EMNLP 2024finding

Recent advancements in multimodal large language models (MLLMs) have made significant progress in integrating information across various modalities, yet real-world applications in educational and scientific domains remain challenging. This paper introduces the Multimodal Scientific ASR (MS-ASR) task…

2024

IMO: Greedy Layer-Wise Sparse Representation Learning for Out-of-Distribution Text Classification with Pre-trained Models

ACL 2024long

Machine learning models have made incredible progress, but they still struggle when applied to examples from unseen domains. This study focuses on a specific problem of domain generalization, where a model is trained on one source domain and tested on multiple target domains that are unseen during t…

2024

MTP: A Dataset for Multi-Modal Turning Points in Casual Conversations

ACL 2024short

Detecting critical moments, such as emotional outbursts or changes in decisions during conversations, is crucial for understanding shifts in human behavior and their consequences. Our work introduces a novel problem setting focusing on these moments as turning points (TPs), accompanied by a meticulo…

Cited by 1SourcePDFScholar
2024

Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models

EMNLP 2024main

Large language models (LLMs) are typically fine-tuned on diverse and extensive datasets sourced from various origins to develop a comprehensive range of skills, such as writing, reasoning, chatting, coding, and more. Each skill has unique characteristics, and these datasets are often heterogeneous a…

Cited by 2SourcePDFScholar
2024

RENOVI: A Benchmark Towards Remediating Norm Violations in Socio-Cultural Conversations

NAACL 2024findings

Norm violations occur when individuals fail to conform to culturally accepted behaviors, which may lead to potential conflicts. Remediating norm violations requires social awareness and cultural sensitivity of the nuances at play. To equip interactive AI systems with a remediation ability, we offer…

2024

Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights

EMNLP 2024main

Large Multimodal Models (LMMs) have achieved great success recently, demonstrating a strong capability to understand multimodal information and to interact with human users. Despite the progress made, the challenge of detecting high-risk interactions in multimodal settings, and in particular in spee…

2022

Fire Burns, Sword Cuts: Commonsense Inductive Bias for Exploration in Text-based Games

ACL 2022short

Text-based games (TGs) are exciting testbeds for developing deep reinforcement learning techniques due to their partially observed environments and large action spaces. In these games, the agent learns to explore the environment via natural language interactions with the game simulator. A fundamenta…