ICASSP 2024accepted0 citations

Fusing Modality-Specific Representations and Decisions for Multimodal Emotion Recognition

Yu-Ping Ruan, Shoukang Han, Taihao Li, Yanfeng Wu

Abstract

Multimodal emotion recognition (MER) is important for building humanoid chatbots and has gained increasing attention in recent years. Existing studies have proven that extracting better modality-specific representations, which keep both commonality and individuality information of different modalities, is important for the MER task. However, all these works are restricted in making final predictions based on fusing modality-specific representations, and the effectiveness of the modality-specific decisions has not been studied. In this paper, we propose for the first time to fuse both the modality-specific representations and decisions for the MER task and design a bi-channel fusing network (BCFN). Specifically, a BCFN model first extracts and mixes the modality-specific representations and decisions in two convolutional blocks respectively, and then fuses the two joint multimodal features for the final decision. Extensive experiments are conducted on two MER benchmark datasets with state-of-the-art models employed as baselines for comparison. The empirical results demonstrate the superiority of our proposed BCFN model and confirm the effectiveness of incorporating modality-specific decisions for the MER task.

BibTeX
@inproceedings{icassp2024_fusingmodalitysp,
  title = {Fusing Modality-Specific Representations and Decisions for Multimodal Emotion Recognition},
  author = {Yu-Ping Ruan and Shoukang Han and Taihao Li and Yanfeng Wu},
  booktitle = {ICASSP 2024},
  year = {2024}
}