2024
Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
ECCV 2024poster
"Recent video-text foundation models have demonstrated strong performance on a wide variety of downstream video understanding tasks. Can these video-text models genuinely understand the contents of natural videos? Standard video-text evaluations could be misleading as many questions can be inferred…