ICASSP 2019accepted0 citations

Implicit Fusion by Joint Audiovisual Training for Emotion Recognition in Mono Modality

Jing Han, Zixing Zhang, Zhao Ren, Björn W. Schuller

Abstract

Despite significant advances in emotion recognition from one individual modality, previous studies fail to take advantage of other modalities to train models in mono-modal scenarios. In this work, we propose a novel joint training model which implicitly fuses audio and visual information in the training procedure for either speech or facial emotion recognition. Specifically, the model consists of one modality-specific network per individual modality and one shared network to map both audio and visual cues into final predictions. In the training process, we additionally take the loss from one auxiliary modality into account besides the main modality. To evaluate the effectiveness of the implicit fusion model, we conduct extensive experiments for mono-modal emotion classification and regression, and find that the implicit fusion models outperform the standard mono-modal training process.

BibTeX
@inproceedings{icassp2019_implicitfusionby,
  title = {Implicit Fusion by Joint Audiovisual Training for Emotion Recognition in Mono Modality},
  author = {Jing Han and Zixing Zhang and Zhao Ren and Björn W. Schuller},
  booktitle = {ICASSP 2019},
  year = {2019}
}
Implicit Fusion by Joint Audiovisual Training for Emotion Recognition in Mono Modality · ICASSP 2019