Multimodal Emotion Recognition with Target Speaker-Based Facial Embeddings
Serin Heo, Jehyun Kyung, Joon-Hyuk Chang
Abstract
Effectively recognizing emotions requires sophisticated approaches for interpreting diverse modalities, particularly in real-world scenarios where multiple data sources, such as speech, text, and visual cues, are often noisy and incomplete. This study proposes an advanced multimodal emotion recognition system that integrates these three modalities by adding the speaker detection and extraction algorithm within visual data. The pre-trained Q-Former used in the proposed system then captures and interprets visual signals supported with designated prompts, resulting in facial-related features that significantly improve emotion recognition performance. We then utilize a cross-modal transformer to unify the visual, speech, and text embeddings for accurate emotion classification. We achieved a 2.9% and 3.3% improvement in accuracy and F1 score, respectively, on the MELD dataset compared to the baseline.
BibTeX
@inproceedings{icassp2025_multimodalemotio,
title = {Multimodal Emotion Recognition with Target Speaker-Based Facial Embeddings},
author = {Serin Heo and Jehyun Kyung and Joon-Hyuk Chang},
booktitle = {ICASSP 2025},
year = {2025}
}