EMNLP 2022finding12 citations

Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis

Iek-Heng Chu, Ziyi Chen, Xinlu Yu, Mei Han, Jing Xiao, Peng Chang

Abstract

Multimodal speech emotion recognition (SER) and sentiment analysis (SA) are important techniques for human-computer interaction. Most existing multimodal approaches utilize either shallow cross-modal fusion of pretrained features, or deep cross-modal fusion with raw features. Recently, attempts have been made to fuse pretrained feature representations in a deep fusion manner during fine-tuning stage. However those approaches have not led to improved results, partially due to their relatively simple fusion mechanisms and lack of proper cross-modal pretraining. In this work, leveraging single-modal pretrained models (RoBERTa and HuBERT), we propose a novel deeply-fused audio-text bi-modal transformer with carefully designed cross-modal fusion mechanism and a stage-wise cross-modal pretraining scheme to fully facilitate the cross-modal learning. Our experiment results show that the proposed method achieves state-of-the-art results on the public IEMOCAP emotion and CMU-MOSEI sentiment datasets, exceeding the previous benchmarks by a large margin.

BibTeX
@inproceedings{chu-etal-2022-self,
    title = "Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis",
    author = "Chu, Iek-Heng  and
      Chen, Ziyi  and
      Yu, Xinlu  and
      Han, Mei  and
      Xiao, Jing  and
      Chang, Peng",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2022",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.findings-emnlp.375/",
    doi = "10.18653/v1/2022.findings-emnlp.375",
    pages = "5105--5114"
}
Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis · EMNLP 2022