Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked visual modeling help mitigate this issue, they still suffer from two fundamental limitations: severe visual information lo