NeurIPS 2021poster218 citations

Predicting What You Already Know Helps: Provable Self-Supervised Learning

Jason D. Lee, Qi Lei, Nikunj Saunshi, Jiacheng Zhuo

Abstract

Self-supervised representation learning solves auxiliary prediction tasks (known as pretext tasks), that do not require labeled data, to learn semantic representations. These pretext tasks are created solely using the input features, such as predicting a missing image patch, recovering the color channels of an image from context, or predicting missing words, yet predicting this \textit{known} information helps in learning representations effective for downstream prediction tasks. This paper posits a mechanism based on approximate conditional independence to formalize how solving certain pretext tasks can learn representations that provably decrease the sample complexity of downstream supervised tasks. Formally, we quantify how the approximate independence between the components of the pretext task (conditional on the label and latent variables) allows us to learn representations that can solve the downstream task with drastically reduced sample complexity by just training a linear layer on top of the learned representation.

statistical learning theoryapproximate conditional independencereconstruction-based self-supervised learning
BibTeX
@inproceedings{
lee2021predicting,
title={Predicting What You Already Know Helps: Provable Self-Supervised Learning},
author={Jason D. Lee and Qi Lei and Nikunj Saunshi and Jiacheng Zhuo},
booktitle={Advances in Neural Information Processing Systems},
editor={A. Beygelzimer and Y. Dauphin and P. Liang and J. Wortman Vaughan},
year={2021},
url={https://openreview.net/forum?id=Yx1OzVU_SRi}
}
Predicting What You Already Know Helps: Provable Self-Supervised Learning · NeurIPS 2021