Trust Your Partner's Friends: Hierarchical Cross-Modal Contrastive Pre-Training for Video-Text Retrieval
Video-text retrieval has greatly benefited from the massive web video in recent years, while the performance is still limited to the weak supervision from the uncurated data. In this work, we propose to leverage the well-represented information of each original modality and exploit complementary inf…