← Search

Mohammad Soleymani

21 accepted papers

2026

AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization

ICLR 2026poster

Emotion understanding is essential for building socially intelligent agents. Although recent multimodal large language models (MLLMs) have shown strong performance on this task, two key challenges remain: (i) spurious associations between emotions and irrelevant audiovisual cues and (ii) hallucinati…

Cited by 0SourcecodeScholar
2026

Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox

ICML 2026poster

Audio large language models (Audio LLMs) demonstrate strong performance on speech understanding tasks, yet their ability to understand paralinguistic information remains limited. To systematically quantify this issue, we introduce VoxParadox, an adversarial benchmark with 2,000 verified examples, sp…

Cited by 0SourceScholar
2026

MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization

CVPR 2026

Omni-modal large language models (omni LLMs) have recently achieved strong performance across audiovisual understanding tasks, yet they remain highly susceptible to cross-modal hallucinations arising from spurious correlations and dominant language priors. In this work, we propose Modality-Decoupled

Cited by 0SourceScholar
2026

Riemannian optimization on the manifold of unitary and symmetric matrices with application to BD-RIS-assisted systems

ICASSP 2026poster

In this paper, we rigorously characterize for the first time the manifold of unitary and symmetric matrices, deriving its tangent space and its geodesics. The resulting parameterization of the geodesics (through a real and symmetric matrix) allows us to derive a new Riemannian manifold optimization…

Cited by 0SourcePDFScholar
2025

DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion

ICCV 2025poster

Generating naturalistic and nuanced listener motions for extended interactions remains an open problem. Existing methods often rely on low-dimensional motion codes for facial behavior generation followed by photorealistic rendering, limiting both visual fidelity and expressive richness. To address t…

2025

X-Dyna: Expressive Dynamic Human Image Animation

CVPR 2025highlight

We introduce X-Dyna, a novel zero-shot, diffusion-based pipeline for animating a single human image using facial expressions and body movements derived from a driving video, that generates realistic, context-aware dynamics for both the subject and the surrounding environment. Building on prior appro…

2024

Build Your Own Robot Friend: An Open-Source Learning Module for Accessible and Engaging AI Education

AAAI 2024technical

As artificial intelligence (AI) is playing an increasingly important role in our society and global economy, AI education and literacy have become necessary components in college and K-12 education to prepare students for an AI-powered society. However, current AI curricula have not yet been made ac…

Cited by 9SourcePDFScholar
2024

DIM: Dyadic Interaction Modeling for Social Behavior Generation

ECCV 2024poster

"Human-human communication is like a delicate dance where listeners and speakers concurrently interact to maintain conversational dynamics. Hence, an effective model for generating listener nonverbal behaviors requires understanding the dyadic context and interaction. In this paper, we present an ef…

Cited by 7SourcePDFScholar
2024

Ex2Eg-MAE: A Framework for Adaptation of Exocentric Video Masked Autoencoders for Egocentric Social Role Understanding

ECCV 2024poster

"Self-supervised learning methods have demonstrated impressive performance across visual understanding tasks, including human behavior understanding. However, there has been limited work for self-supervised learning for egocentric social videos. Visual processing in such contexts faces several chall…

Cited by 1SourcePDFScholar
2024

MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion

ICML 2024poster

In this work, we propose MagicPose, a diffusion-based model for 2D human pose and facial expression retargeting. Specifically, given a reference image, we aim to generate a person's new images by controlling the poses and facial expressions while keeping the identity unchanged. To this end, we propo…

2023

Interference Leakage Minimization in RIS-Assisted MIMO Interference Channels

ICASSP 2023accepted

We address the problem of interference leakage (IL) minimization in the K-user multiple-input multiple-output (MIMO) interference channel (IC) assisted by a reconfigurable intelligent surface (RIS). We describe an iterative algorithm based on block coordinate descent to minimize the IL cost function…

Cited by 0SourceScholar
2022

Self-Supervised Learning for Sentiment Analysis via Image-Text Matching

ICASSP 2022accepted

There is often a resemblance in the sentiment expressed in social media posts (text) and their accompanying images. In this paper, We leverage this sentiment congruence for self-supervised representation learning for sentiment analysis. By teaching the model to pair an image with its corresponding s…

Cited by 0SourceScholar
2021

Multimodal Phased Transformer for Sentiment Analysis

EMNLP 2021main

Multimodal Transformers achieve superior performance in multimodal learning tasks. However, the quadratic complexity of the self-attention mechanism in Transformers limits their deployment in low-resource devices and makes their inference and training computationally expensive. We propose multimodal…

2021

Speaker Turn Modeling for Dialogue Act Classification

EMNLP 2021finding

Dialogue Act (DA) classification is the task of classifying utterances with respect to the function they serve in a dialogue. Existing approaches to DA classification model utterances without incorporating the turn changes among speakers throughout the dialogue, therefore treating it no different th…

2021

Subject-Invariant Eeg Representation Learning For Emotion Recognition

ICASSP 2021accepted

The discrepancies between the distributions of the train and test data, a.k.a., domain shift, result in lower generalization for emotion recognition methods. One of the main factors contributing to these discrepancies is human variability. Domain adaptation methods are developed to alleviate the pro…

Cited by 0SourceScholar
2020

Expression-Guided EEG Representation Learning for Emotion Recognition

ICASSP 2020accepted

Learning a joint and coordinated representation between different modalities can improve multimodal emotion recognition. In this paper, we propose a deep representation learning approach for emotion recognition from electroencephalogram (EEG) signals guided by facial electromyogram (EMG) and electro…

Cited by 0SourceScholar
2020

Towards A Friendly Online Community: An Unsupervised Style Transfer Framework for Profanity Redaction

COLING 2020main

Offensive and abusive language is a pressing problem on social media platforms. In this work, we propose a method for transforming offensive comments, statements containing profanity or offensive language, into non-offensive ones. We design a Retrieve, Generate and Edit unsupervised style transfer p…

Cited by 35SourcePDFScholar
2019

Energy-efficient Design for Underlay Cognitive Radio Using Improper Signaling

ICASSP 2019accepted

Improper Gaussian signaling (IGS) has been used as an effective interference management tool in interference limited systems. Improper Gaussian signals are correlated with their complex conjugates. In this paper, we investigate the optimality of IGS from an energy efficiency (EE) perspective. First,…

Cited by 0SourceScholar