ICASSP 2024accepted0 citations

End-To-End Spatially-Constrained Multi-Perspective Fine-Grained Image Captioning

Yifan Zhang, Chunzhen Lin, Donglin Cao, Dazhen Lin

Abstract

The perspective of captions in fine-grained image captioning crucially impacts people’s perception and understanding of the image. However, existing methods often overlook this aspect, resulting in captions that struggle to accurately convey the image’s hierarchical and spatial information. In this paper, we propose an end-to-end Spatially-Constrained multi-perspective fine-grained Image Captioning (SCIC) model. SCIC initially predicts the optimal perspective for captioning the image and subsequently utilizes this optimal perspective as a constraint condition for caption generation. Furthermore, SCIC is capable of generating multi-perspective image captions based on customized perspectives. Experimental results show that our model effectively improves the state-of-the-art CIDEr score by about 14.38% and can generate multi-perspective fine-grained captions for the same image.

BibTeX
@inproceedings{icassp2024_endtoendspatiall,
  title = {End-To-End Spatially-Constrained Multi-Perspective Fine-Grained Image Captioning},
  author = {Yifan Zhang and Chunzhen Lin and Donglin Cao and Dazhen Lin},
  booktitle = {ICASSP 2024},
  year = {2024}
}
End-To-End Spatially-Constrained Multi-Perspective Fine-Grained Image Captioning · ICASSP 2024