AAAI 2023technical10 citations

Video-Audio Domain Generalization via Confounder Disentanglement

Shengyu Zhang, Xusheng Feng, Wenyan Fan, Wenjing Fang, Fuli Feng, Wei Ji, Shuo Li, Li Wang

Abstract

Existing video-audio understanding models are trained and evaluated in an intra-domain setting, facing performance degeneration in real-world applications where multiple domains and distribution shifts naturally exist. The key to video-audio domain generalization (VADG) lies in alleviating spurious correlations over multi-modal features. To achieve this goal, we resort to causal theory and attribute such correlation to confounders affecting both video-audio features and labels. We propose a DeVADG framework that conducts uni-modal and cross-modal deconfounding through back-door adjustment. DeVADG performs cross-modal disentanglement and obtains fine-grained confounders at both class-level and domain-level using half-sibling regression and unpaired domain transformation, which essentially identifies domain-variant factors and class-shared factors that cause spurious correlations between features and false labels. To promote VADG research, we collect a VADG-Action dataset for video-audio action recognition with over 5,000 video clips across four domains (e.g., cartoon and game) and ten action classes (e.g., cooking and riding). We conduct extensive experiments, i.e., multi-source DG, single-source DG, and qualitative analysis, validating the rationality of our causal analysis and the effectiveness of the DeVADG framework.

BibTeX
@article{Zhang_Feng_Fan_Fang_Feng_Ji_Li_Wang_Zhao_Zhao_Chua_Wu_2023, title={Video-Audio Domain Generalization via Confounder Disentanglement}, volume={37}, url={https://ojs.aaai.org/index.php/AAAI/article/view/26787}, DOI={10.1609/aaai.v37i12.26787}, abstractNote={Existing video-audio understanding models are trained and evaluated in an intra-domain setting, facing performance degeneration in real-world applications where multiple domains and distribution shifts naturally exist. The key to video-audio domain generalization (VADG) lies in alleviating spurious correlations over multi-modal features. To achieve this goal, we resort to causal theory and attribute such correlation to confounders affecting both video-audio features and labels. We propose a DeVADG framework that conducts uni-modal and cross-modal deconfounding through back-door adjustment. DeVADG performs cross-modal disentanglement and obtains fine-grained confounders at both class-level and domain-level using half-sibling regression and unpaired domain transformation, which essentially identifies domain-variant factors and class-shared factors that cause spurious correlations between features and false labels. To promote VADG research, we collect a VADG-Action dataset for video-audio action recognition with over 5,000 video clips across four domains (e.g., cartoon and game) and ten action classes (e.g., cooking and riding). We conduct extensive experiments, i.e., multi-source DG, single-source DG, and qualitative analysis, validating the rationality of our causal analysis and the effectiveness of the DeVADG framework.}, number={12}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Zhang, Shengyu and Feng, Xusheng and Fan, Wenyan and Fang, Wenjing and Feng, Fuli and Ji, Wei and Li, Shuo and Wang, Li and Zhao, Shanshan and Zhao, Zhou and Chua, Tat-Seng and Wu, Fei}, year={2023}, month={Jun.}, pages={15322-15330} }
Video-Audio Domain Generalization via Confounder Disentanglement · AAAI 2023