Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-Training
In this work we tackle the problem of unsupervised domain adaptation (UDA) for video action recognition. Our approach which we call UNITE uses an image teacher model to adapt a video student model to the target domain. UNITE first employs self-supervised pre-training to promote discriminative featur…