2025
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
AAAI 2025technical
Weakly-Supervised Dense Video Captioning (WSDVC) aims to localize and describe all events of interest in a video without requiring annotations of event boundaries. This setting poses a great challenge in accurately locating the temporal location of event, as the relevant supervision is unavailable.…