Interventional Contrastive Learning with Meta Semantic Regularizer
Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Bing Su, Hui Xiong
Abstract
Contrastive learning (CL)-based self-supervised learning models learn visual representations in a pairwise manner. Although the prevailing CL model has achieved great progress, in this paper, we uncover an ever-overlooked phenomenon: When the CL model is trained with full images, the performance tested in full images is better than that in foreground areas; when the CL model is trained with foreground areas, the performance tested in full images is worse than that in foreground areas. This observation reveals that backgrounds in images may interfere with the model learning semantic information and their influence has not been fully eliminated. To tackle this issue, we build a Structural Causal Model (SCM) to model the background as a confounder. We propose a backdoor adjustment-based regularization method, namely
BibTeX
@InProceedings{pmlr-v162-qiang22a,
title = {Interventional Contrastive Learning with Meta Semantic Regularizer},
author = {Qiang, Wenwen and Li, Jiangmeng and Zheng, Changwen and Su, Bing and Xiong, Hui},
booktitle = {Proceedings of the 39th International Conference on Machine Learning},
pages = {18018--18030},
year = {2022},
editor = {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
volume = {162},
series = {Proceedings of Machine Learning Research},
month = {17--23 Jul},
publisher = {PMLR},
pdf = {https://proceedings.mlr.press/v162/qiang22a/qiang22a.pdf},
url = {https://proceedings.mlr.press/v162/qiang22a.html},
abstract = {Contrastive learning (CL)-based self-supervised learning models learn visual representations in a pairwise manner. Although the prevailing CL model has achieved great progress, in this paper, we uncover an ever-overlooked phenomenon: When the CL model is trained with full images, the performance tested in full images is better than that in foreground areas; when the CL model is trained with foreground areas, the performance tested in full images is worse than that in foreground areas. This observation reveals that backgrounds in images may interfere with the model learning semantic information and their influence has not been fully eliminated. To tackle this issue, we build a Structural Causal Model (SCM) to model the background as a confounder. We propose a backdoor adjustment-based regularization method, namely