Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions
In recent text-video retrieval, the use of additional captions from vision-language models has shown promising effects on the performance. However, existing models using additional captions often have struggled to capture the rich semantics, including temporal changes, inherent in the video. In addi…