ACL 2024findings14 citations

SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval

Siwei Wu, Yizhi Li, Kang Zhu, Ge Zhang, Yiming Liang, Kaijing Ma, Chenghao Xiao, Haoran Zhang

Abstract

Multi-modal information retrieval (MMIR) is a rapidly evolving field where significant progress has been made through advanced representation learning and cross-modality alignment research, particularly in image-text pairing.However, current benchmarks for evaluating MMIR performance on image-text pairings overlook the scientific domain, which has a notable gap with the generic data since the caption of scientific charts and tables usually describes the analysis of experimental results or scientific principles in contrast to human activity or scenery depicted in generic images.To bridge this gap, we develop a scientific domain-specific MMIR benchmark (SciMMIR) by leveraging open-access research paper corpora to extract data relevant to the scientific domain. This benchmark comprises 530K meticulously curated image-text pairs, extracted from figures and tables with detailed captions from scientific documents.We further annotate the image-text pairs with a two-level subset-subcategory hierarchy to facilitate a more comprehensive evaluation of the baselines. We conduct zero-shot and fine-tuned evaluations on prominent multi-modal image-captioning and visual language models, such as CLIP, BLIP, and BLIP-2.Our findings offer critical insights for MMIR in the scientific domain, including the impact of pre-training and fine-tuning settings and the effects of different visual and textual encoders.

BibTeX
@inproceedings{wu-etal-2024-scimmir,
    title = "{S}ci{MMIR}: Benchmarking Scientific Multi-modal Information Retrieval",
    author = "Wu, Siwei  and
      Li, Yizhi  and
      Zhu, Kang  and
      Zhang, Ge  and
      Liang, Yiming  and
      Ma, Kaijing  and
      Xiao, Chenghao  and
      Zhang, Haoran  and
      Yang, Bohao  and
      Chen, Wenhu  and
      Huang, Wenhao  and
      Al Moubayed, Noura  and
      Fu, Jie  and
      Lin, Chenghua",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2024",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.findings-acl.746/",
    doi = "10.18653/v1/2024.findings-acl.746",
    pages = "12560--12574"
}
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval · ACL 2024