ASCNet: Self-Supervised Video Representation Learning With Appearance-Speed Consistency
We study self-supervised video representation learning, which is a challenging task due to 1) sufficient labels for supervision; 2) unstructured and noisy visual information. Existing methods mainly use contrastive loss with video clips as the instances and learn visual representation by discriminat…