DEVIAS: Learning Disentangled Video Representations of Action and Scene
"Video recognition models often learn scene-biased action representation due to the spurious correlation between actions and scenes in the training data. Such models show poor performance when the test data consists of videos with unseen action-scene combinations. Although Scene-debiased action reco…