ICASSP 2025accepted0 citations

Video-Poetry Retrieval with Multimodal Knowledge Graph Guided Unsupervised Pre-training

Xinru Wei, Yuqing Li, Bin Wu

Abstract

Classical Chinese poetry, with its rich cultural heritage, holds immense artistic value. Recently, research on multi-modal approaches to classical poetry has garnered attention. Current research mainly focuses on poetry and images, but images cannot fully capture dynamic scenes like sunrises. However, the inherent temporal nature of video can vividly present these scene transitions. In this paper, we introduce a novel task: video-poetry retrieval. Due to the disparities between classical Chinese poetry and modern Chinese, and the time-consuming nature of traditional supervised methods for collecting parallel data. We propose VPR-MLCL, a video-poetry pre-training model based on unsupervised pre-training methods from vision-language field to address these issues. VPR-MLCL uses a multi-modal knowledge graph of classical Chinese poetry to bridge the gap, learning cross-modal alignment from video-caption pairs and poetry-image entities rather than direct video-poetry pairs. Additionally, we have constructed a video-poetry dataset for this task, and experiments demonstrate the effectiveness of VPR-MLCL.

BibTeX
@inproceedings{icassp2025_videopoetryretri,
  title = {Video-Poetry Retrieval with Multimodal Knowledge Graph Guided Unsupervised Pre-training},
  author = {Xinru Wei and Yuqing Li and Bin Wu},
  booktitle = {ICASSP 2025},
  year = {2025}
}