Text-guided Multimodal Fusion for the Multimodal Emotion and Intent Joint Understanding
Yu Zhang, Bin Chen, Hongfei Ye, Zijian Gao, Tianjiao Wan, Long Lan, Kele Xu
Abstract
Emotion and Intent Joint Understanding in Multi-modal Conversation is a challenging task in the field of affective computing, aiming to decode the semantic information manifested in the multimodal conversational while simultaneously inferring the emotions and intents of the utterance. To address this challenge, we propose the Text-guided Multimodal Emotion-Intent Joint Recognition method. By leveraging the text modality to guide the fusion process, it effectively reduces the noise introduced by other modalities. To strengthen the text modality’s guiding role, we use large language models (LLMs) for multi-turn targeted data augmentation and oversampling strategies to address data imbalance. Our approach achieved first place in Track 1 (English) of the ICASSP 2025 MEIJU Challenge, demonstrating its effectiveness in practical applications.
BibTeX
@inproceedings{icassp2025_textguidedmultim,
title = {Text-guided Multimodal Fusion for the Multimodal Emotion and Intent Joint Understanding},
author = {Yu Zhang and Bin Chen and Hongfei Ye and Zijian Gao and Tianjiao Wan and Long Lan and Kele Xu},
booktitle = {ICASSP 2025},
year = {2025}
}