2025
VideoGEM: Training-free Action Grounding in Videos
CVPR 2025poster
Vision-language foundation models have shown impressive capabilities across various zero-shot tasks, including training-free localization and grounding, primarily focusing on localizing objects in images. However, leveraging those capabilities to localize actions and events in videos is challenging,…