ICASSP 2025accepted0 citations

GateM2Former: Gated Feature Selection and Expert Modeling in Multimodal Emotion Recognition

Weixiang Xu, Zhongren Dong, Runming Wang, Xinzhou Xu, Zixing Zhang

Abstract

In recent years, multimodal emotion recognition (MER) has gained significant attention due to its potential to integrate information from diverse signals. However, existing methods often struggle to effectively capture complex interactions and contextual information both inter- and intra-modalities, and even to extract the salient representations from pre-trained models. To address these issues, we propose a novel model, gated Mixture of Multimodal Experts (MoME) and Mixtral of Experts (MixMoE) models, namely GateM<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup>Former. The gate mechanism is used to select the most relevant representations from pre-trained models. The MoME and MixMoE expert modules respectively learn the individual characteristics of each modality and the intrinsic alignment and potential interactions between modalities. Besides, we design a hierarchical merge structure to better suit the long sequence scenario (i. e., speech in our case). To verify the effectiveness of the introduced model, we conducted extensive experiments on the IEMOCAP and MELD datasets. The results show that GateM<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup>Former, with a universal multimodal structure, is able to achieve the best results on IEMOCAP and MELD compared with other latest approaches.

BibTeX
@inproceedings{icassp2025_gatem2formergate,
  title = {GateM2Former: Gated Feature Selection and Expert Modeling in Multimodal Emotion Recognition},
  author = {Weixiang Xu and Zhongren Dong and Runming Wang and Xinzhou Xu and Zixing Zhang},
  booktitle = {ICASSP 2025},
  year = {2025}
}