← Search

Sai Zhou

1 accepted papers

2024

Align Before Adapt: Leveraging Entity-to-Region Alignments for Generalizable Video Action Recognition

CVPR 2024poster

Large-scale visual-language pre-trained models have achieved significant success in various video tasks. However most existing methods follow an "adapt then align" paradigm which adapts pre-trained image encoders to model video-level representations and utilizes one-hot or text embedding of the acti…

Cited by 9SourcePDFScholar