A Simple Recipe for Contrastively Pre-training Video-First Encoders Beyond 16 Frames
Understanding long real-world videos requires modeling of long-range visual dependencies. To this end we explore video-first architectures building on the common paradigm of transferring large-scale image--text models to video via shallow temporal fusion. However we expose two limitations to the app…