AAAI 2023technical3 citations

Learning Semantic Alignment with Global Modality Reconstruction for Video-Language Pre-training towards Retrieval

Mingchao Li, Xiaoming Shi, Haitao Leng, Wei Zhou, Hai-Tao Zheng, Kuncai Zhang

Abstract

Video-language pre-training for text-based video retrieval tasks is vitally important. Previous pre-training methods suffer from the semantic misalignments. The reason is that these methods ignore sequence alignments but focusing on critical token alignment. To alleviate the problem, we propose a video-language pre-training framework, termed videolanguage pre-training For lEarning sEmantic aLignments (FEEL), to learn semantic alignments at the sequence level. Specifically, the global modality reconstruction and the cross- modal self-contrasting method is utilized to learn the alignments at the sequence level better. Extensive experimental results demonstrate the effectiveness of FEEL on text-based video retrieval and text-based video corpus moment retrieval.

BibTeX
@article{Li_Shi_Leng_Zhou_Zheng_Zhang_2023, title={Learning Semantic Alignment with Global Modality Reconstruction for Video-Language Pre-training towards Retrieval}, volume={37}, url={https://ojs.aaai.org/index.php/AAAI/article/view/25222}, DOI={10.1609/aaai.v37i1.25222}, abstractNote={Video-language pre-training for text-based video retrieval tasks is vitally important. Previous pre-training methods suffer from the semantic misalignments. The reason is that these methods ignore sequence alignments but focusing on critical token alignment. To alleviate the problem, we propose a video-language pre-training framework, termed videolanguage pre-training For lEarning sEmantic aLignments (FEEL), to learn semantic alignments at the sequence level. Specifically, the global modality reconstruction and the cross- modal self-contrasting method is utilized to learn the alignments at the sequence level better. Extensive experimental results demonstrate the effectiveness of FEEL on text-based video retrieval and text-based video corpus moment retrieval.}, number={1}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Li, Mingchao and Shi, Xiaoming and Leng, Haitao and Zhou, Wei and Zheng, Hai-Tao and Zhang, Kuncai}, year={2023}, month={Jun.}, pages={1377-1385} }