Caption Unification for Multi-View Lifelogging Images Based on In-Context Learning with Heterogeneous Semantic Contents
Masaya Sato, Keisuke Maeda, Ren Togo, Takahiro Ogawa, Miki Haseyama
Abstract
This paper presents a new task of caption unification and a novel caption unification method for multi-view lifelogging images based on in-context learning with heterogeneous semantic contents. Most of the existing image captioning models target a single image and do not consider the common semantic contents among multiple images. Therefore, they suffer from inconsistent captioning for multi-view lifelogging images taken in the same scene, i.e., the images that have common semantic contents regardless of the viewpoints. The proposed method enables training the common semantic contents of multiple images and generating captions that comprehensively contain the semantic contents of the images based on in-context learning. The proposed method has the following two contributions. First, we identify the problem of existing image captioning models and consider caption unification as a new task for multi-view lifelogging images. Second, we generate unified captions based on in-context learning with heterogeneous semantic contents and evaluate the captions from two perspectives, uniformity and retainability. The experimental results demonstrate that the proposed method realizes successful caption unification.
BibTeX
@inproceedings{icassp2024_captionunificati,
title = {Caption Unification for Multi-View Lifelogging Images Based on In-Context Learning with Heterogeneous Semantic Contents},
author = {Masaya Sato and Keisuke Maeda and Ren Togo and Takahiro Ogawa and Miki Haseyama},
booktitle = {ICASSP 2024},
year = {2024}
}