ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders
"We propose , a model that combines both Masked AutoEncoders (MAE) and contrastive learning. is trained using a global representation obtained by pooling the local features learned under an MAE reconstruction loss and using this representation under a contrastive objective across images and video fr…