Video-Poetry Retrieval with Multimodal Knowledge Graph Guided Unsupervised Pre-training
Abstract
Classical Chinese poetry, with its rich cultural heritage, holds immense artistic value. Recently, research on multi-modal approaches to classical poetry has garnered attention. Current research mainly focuses on poetry and images, but images cannot fully capture dynamic scenes like sunrises. However, the inherent temporal nature of video can vividly present these scene transitions. In this paper, we introduce a novel task: video-poetry retrieval. Due to the disparities between classical Chinese poetry and modern Chinese, and the time-consuming nature of traditional supervised methods for collecting parallel data. We propose VPR-MLCL, a video-poetry pre-training model based on unsupervised pre-training methods from vision-language field to address these issues. VPR-MLCL uses a multi-modal knowledge graph of classical Chinese poetry to bridge the gap, learning cross-modal alignment from video-caption pairs and poetry-image entities rather than direct video-poetry pairs. Additionally, we have constructed a video-poetry dataset for this task, and experiments demonstrate the effectiveness of VPR-MLCL.
BibTeX
@inproceedings{icassp2025_videopoetryretri,
title = {Video-Poetry Retrieval with Multimodal Knowledge Graph Guided Unsupervised Pre-training},
author = {Xinru Wei and Yuqing Li and Bin Wu},
booktitle = {ICASSP 2025},
year = {2025}
}