Vid-Group: Temporal Video Grounding Pretraining from Unlabeled Videos in the Wild
Given a natural language query, temporal video grounding aims to localize the described temporal moment in an untrimmed video. A major challenge of this task is its heavy dependence on labor-intensive annotations for training. Unlike existing works that directly train models on manually curated data…