Difference Bonds Consistency and Complementarity to Enhance Multimodal Representation Learning
Congbing He, Sensen Song, Zhenhong Jia, Hui Zhao
Abstract
In the field of multimodal representation learning, existing research has primarily focused on exploring modal consistency and modal complementarity, while overlooking the positive role of modal difference. Moreover, modal difference establishes a bonding relationship between modal consistency and modal complementarity. However, existing algorithms lack the study of this relationship, resulting in an accuracy that still needs to be improved on multimodal classification tasks. To tackle the above issues, We propose a novel multimodal representation learning framework. It enhances multimodal representation learning by modal difference constructing connecting bonds of modal consistency and modal complementarity. We conducted experiments on two widely used multimodal emotion recognition datasets, IEMOCAP and MELD. The results demonstrate that our method outperforms existing multimodal representation learning approaches in terms of accuracy on the multimodal emotion recognition task.
BibTeX
@inproceedings{icassp2025_differencebondsc,
title = {Difference Bonds Consistency and Complementarity to Enhance Multimodal Representation Learning},
author = {Congbing He and Sensen Song and Zhenhong Jia and Hui Zhao},
booktitle = {ICASSP 2025},
year = {2025}
}