ICASSP 2025accepted0 citations

Spatiotemporal-Aware Visual Captioning using Vision-Language Pre-Training Model

Shuai Wu, Weidong Yang, Shuyan Wu

Abstract

Current visual captioning technologies typically transform 3D/2D visual information into one-dimensional sequential data and employ language models to generate corresponding descriptions. This approach, however, compromises the spatiotemporal information in visual data, making it difficult for models to capture temporal variations and the relative spatial relationships between objects. To address this issue, we propose STPos-VC, a pre-trained vision-language model that maps visual information from the visual vector space to the textual vector space through a visual-text mapper and generates natural language descriptions using a decoder. The mapper incorporates three-dimensional rotational position encoding, which effectively preserves the relative spatiotemporal positional relationships. Furthermore, we pre-train the model on a mixed dataset comprising images and videos through a visual question-answering framework, enabling the model to perform well even with small sample sizes. Experimental results across multiple datasets demonstrate that, compared to existing methods, STPos-VC achieves superior performance in both general-purpose and domain-specific applications.

BibTeX
@inproceedings{icassp2025_spatiotemporalaw,
  title = {Spatiotemporal-Aware Visual Captioning using Vision-Language Pre-Training Model},
  author = {Shuai Wu and Weidong Yang and Shuyan Wu},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Spatiotemporal-Aware Visual Captioning using Vision-Language Pre-Training Model · ICASSP 2025