EMNLP 2024finding1 citations

Infrared-LLaVA: Enhancing Understanding of Infrared Images in Multi-Modal Large Language Models

Shixin Jiang, Zerui Chen, Jiafeng Liang, Yanyan Zhao, Ming Liu, Bing Qin

Abstract

Expanding the understanding capabilities of multi-modal large language models (MLLMs) for infrared modality is a challenge due to the single-modality nature and limited amount of training data. Existing methods typically construct a uniform embedding space for cross-modal alignment and leverage abundant visual image data to indirectly understand infrared images. However, they ignore the supervisory signals of infrared-modality-specific attributes, which may lead to biased understanding of infrared images. To address this issue, we propose a debating multi-agent generation system which transfers knowledge from visible images to generate infrared image-text pairs and infrared instruction data. Moreover, we construct an infrared question-answering benchmark based on common infrared tasks. Experimental results from incremental fine-tuning on existing models and our Infrared-LLaVA-7B trained from scratch on infrared data demonstrate the effectiveness of the generated data and the feasibility of the generation approach.

BibTeX
@inproceedings{jiang-etal-2024-infrared,
    title = "Infrared-{LL}a{VA}: Enhancing Understanding of Infrared Images in Multi-Modal Large Language Models",
    author = "Jiang, Shixin  and
      Chen, Zerui  and
      Liang, Jiafeng  and
      Zhao, Yanyan  and
      Liu, Ming  and
      Qin, Bing",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2024",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.findings-emnlp.501/",
    doi = "10.18653/v1/2024.findings-emnlp.501",
    pages = "8573--8591"
}
Infrared-LLaVA: Enhancing Understanding of Infrared Images in Multi-Modal Large Language Models · EMNLP 2024