2025
Kronecker Mask and Interpretive Prompts are Language-Action Video Learners
ICLR 2025poster
Contrastive language-image pretraining (CLIP) has significantly advanced image-based vision learning. A pressing topic subsequently arises: how can we effectively adapt CLIP to the video domain? Recent studies have focused on adjusting either the textual or visual branch of CLIP for action recogniti…