COLING 2025main1 citations

MPID: A Modality-Preserving and Interaction-Driven Fusion Network for Multimodal Sentiment Analysis

Tianyi Li, Daming Liu

Abstract

The advancement of social media has intensified interest in the research direction of Multimodal Sentiment Analysis (MSA). However, current methodologies exhibit relative limitations, particularly in their fusion mechanisms that overlook nuanced differences and similarities across modalities, leading to potential biases in MSA. In addition, indiscriminate fusion across modalities can introduce unnecessary complexity and noise, undermining the effectiveness of the analysis. In this essay, a Modal-Preserving and Interaction-Driven Fusion Network is introduced to address the aforementioned challenges. The compressed representations of each modality are initially obtained through a Token Refinement Module. Subsequently, we employ a Dual Perception Fusion Module to integrate text with audio and a separate Adaptive Graded Fusion Module for text and visual data. The final step leverages text representation to enhance composite representation. Our experiments on CMU-MOSI, CMU-MOSEI, and CH-SIMS datasets demonstrate that our model achieves state-of-the-art performance.

BibTeX
@inproceedings{li-liu-2025-mpid,
    title = "{MPID}: A Modality-Preserving and Interaction-Driven Fusion Network for Multimodal Sentiment Analysis",
    author = "Li, Tianyi  and
      Liu, Daming",
    editor = "Rambow, Owen  and
      Wanner, Leo  and
      Apidianaki, Marianna  and
      Al-Khalifa, Hend  and
      Eugenio, Barbara Di  and
      Schockaert, Steven",
    booktitle = "Proceedings of the 31st International Conference on Computational Linguistics",
    month = jan,
    year = "2025",
    address = "Abu Dhabi, UAE",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.coling-main.291/",
    pages = "4313--4322"
}
MPID: A Modality-Preserving and Interaction-Driven Fusion Network for Multimodal Sentiment Analysis · COLING 2025