← Search

Nixiuming

1 accepted papers

2025

Kronecker Mask and Interpretive Prompts are Language-Action Video Learners

ICLR 2025poster

Contrastive language-image pretraining (CLIP) has significantly advanced image-based vision learning. A pressing topic subsequently arises: how can we effectively adapt CLIP to the video domain? Recent studies have focused on adjusting either the textual or visual branch of CLIP for action recogniti…