Automatic Captioning based on Visible and Infrared Images
Yan Wang, Shuli Lou, Kai Wang, Yunzhe Wang, Xiaohu Yuan, Huaping Liu
Abstract
In this paper, we tackle the task of image captioning with the complementarity of visible light images and infrared images. To address this problem, we propose an RGBIR image fusion captioning model, which can take full advantage of visible light images and infrared images under different conditions. Meanwhile, we develop a wearable environment-assisted system. In addition, we collect and annotate a new dataset containing 3510 pairs of RGB-IR images to support model training. Finally, we conduct extensive experiments to evaluate the model and system. Experimental results show that our new method and system significantly outperform baselines on multiple metrics and have potential practical value.
BibTeX
@inproceedings{icra2024_automaticcaption,
title = {Automatic Captioning based on Visible and Infrared Images},
author = {Yan Wang and Shuli Lou and Kai Wang and Yunzhe Wang and Xiaohu Yuan and Huaping Liu},
booktitle = {ICRA 2024},
year = {2024}
}