ICASSP 2025accepted0 citations

Multimodal Fusion for EEG Emotion Recognition in Music with a Multi-Task Learning Framework

Shengyao Huang, Zhishuo Jin, Dongdong Li, Jinchen Han, Xie Tao

Abstract

This paper proposes a novel EEG-based emotion recognition approach for music, employing a two-stage training framework that integrates emotion representations from music, lyrics, and EEG. First, a modality-specific feature extraction strategy fine-tunes encoders for music and lyrics to extract emotion-related features, while the EEG encoder is fine-tuned to capture identity-related features. The second stage applies a local-to-global fusion strategy, merging EEG and music features for detailed, time-aligned modeling, and integrating global semantic representations from lyrics. Additionally, a multi-task learning framework is utilized to disentangle emotion-related and identity-related information in EEG, enabling robust, subject-independent emotion recognition. Experimental results on the EREMUS dataset show a significant improvement over baseline models, achieving a total score of 51.97%, demonstrating the effectiveness of multimodal integration and our proposed framework for emotion recognition in music.

BibTeX
@inproceedings{icassp2025_multimodalfusion,
  title = {Multimodal Fusion for EEG Emotion Recognition in Music with a Multi-Task Learning Framework},
  author = {Shengyao Huang and Zhishuo Jin and Dongdong Li and Jinchen Han and Xie Tao},
  booktitle = {ICASSP 2025},
  year = {2025}
}