← Search

Gourab Kundu

2 accepted papers

2022

Normalized Contrastive Learning for Text-Video Retrieval

EMNLP 2022main

Cross-modal contrastive learning has led the recent advances in multimodal retrieval with its simplicity and effectiveness. In this work, however, we reveal that cross-modal contrastive learning suffers from incorrect normalization of the sum retrieval probabilities of each text or video instance. S…

Cited by 12SourcePDFScholar
2020

SF-Net: Single-Frame Supervision for Temporal Action Localization

ECCV 2020poster

In this paper, we study an intermediate form of supervision, i.e., single-frame supervision, for temporal action localization (TAL). To obtain the single-frame supervision, the annotators are asked to identify only a single frame within the temporal window of an action. This can significantly reduce…