2024
JM-CLIP: A Joint Modal Similarity Contrastive Learning Model for Video-Text Retrieval
ICASSP 2024accepted
In recent years, the work on video-text retrieval has been well-developed due to the emergence of large-scale pre-training methods. However, these works focus solely on inter-modal interactions and contrasts, neglecting the contrasts of multigrained features within modalities, which makes the simila…