← Search

Yizhen Chen

1 accepted papers

2023

Tagging before Alignment: Integrating Multi-Modal Tags for Video-Text Retrieval

AAAI 2023technical

Vision-language alignment learning for video-text retrieval arouses a lot of attention in recent years. Most of the existing methods either transfer the knowledge of image-text pretraining model to video-text retrieval task without fully exploring the multi-modal information of videos, or simply fus…

Cited by 25SourcePDFScholar