COLING 2024main8 citations

Autonomous Aspect-Image Instruction a2II: Q-Former Guided Multimodal Sentiment Classification

Junjia Feng, Mingqian Lin, Lin Shang, Xiaoying Gao

Abstract

Multimodal aspect-oriented sentiment classification (MABSC) task has garnered significant attention, which aims to identify the sentiment polarities of aspects by combining both language and vision information. However, the limited multimodal data in this task has become a big gap for the vision-language multimodal fusion. While large-scale vision-language pretrained models have been adapted to multiple tasks, their use for MABSC task is still in a nascent stage. In this work, we present an attempt to use the instruction tuning paradigm to MABSC task and leverage the ability of large vision-language models to alleviate the limitation in the fusion of textual and image modalities. To tackle the problem of potential irrelevance between aspects and images, we propose a plug-and-play selector to autonomously choose the most appropriate instruction from the instruction pool, thereby reducing the impact of irrelevant image noise on the final sentiment classification results. We conduct extensive experiments in various scenarios and our model achieves state-of-the-art performance on benchmark datasets, as well as in few-shot settings.

BibTeX
@inproceedings{feng-etal-2024-autonomous,
    title = "Autonomous Aspect-Image Instruction a2{II}: {Q}-Former Guided Multimodal Sentiment Classification",
    author = "Feng, Junjia  and
      Lin, Mingqian  and
      Shang, Lin  and
      Gao, Xiaoying",
    editor = "Calzolari, Nicoletta  and
      Kan, Min-Yen  and
      Hoste, Veronique  and
      Lenci, Alessandro  and
      Sakti, Sakriani  and
      Xue, Nianwen",
    booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "https://aclanthology.org/2024.lrec-main.180/",
    pages = "1996--2005"
}
Autonomous Aspect-Image Instruction a2II: Q-Former Guided Multimodal Sentiment Classification · COLING 2024