CLIP-MSA: Incorporating Inter-Modal Dynamics and Common Knowledge to Multimodal Sentiment Analysis With Clip
Qi Huang, Pingting Cai, Tanyue Nie, Jinshan Zeng
Abstract
Multimodal Sentiment Analysis (MSA) aims to yield the sentiment polarities of speakers in video streams based on multiple modal features such as textual, acoustic and visual features, and has attracted amounts of attention in recent years. Existing MSA models often yield unimodal embeddings from the associated modal features individually, while overlooking the importance of inter-modal dynamics and common knowledge in the extraction of unimodal embeddings, resulting in the limited performance. In this paper, we suggest a novel MSA model called CLIP-MSA through incorporating the inter-modal dynamics and common knowledge into the generation of unimodal representations with the Contrastive Language-Image Pre-training (CLIP), and fusing the textual, acoustic and visual representations with a hierarchical co-attention mechanism. Numerous experimental results over two benchmark datasets show that the proposed model outperforms existing state-of-the-art models on CMU-MOSI, and provides competitive performance on CMU-MOSEI, in terms of four commonly used evaluation metrics.
BibTeX
@inproceedings{icassp2024_clipmsaincorpora,
title = {CLIP-MSA: Incorporating Inter-Modal Dynamics and Common Knowledge to Multimodal Sentiment Analysis With Clip},
author = {Qi Huang and Pingting Cai and Tanyue Nie and Jinshan Zeng},
booktitle = {ICASSP 2024},
year = {2024}
}