A Survey on Masked Autoencoder for Visual Self-supervised Learning
Chaoning Zhang, Chenshuang Zhang, Junha Song, John Seon Keun Yi, In So Kweon
Abstract
With the increasing popularity of masked autoencoders, self-supervised learning (SSL) in vision undertakes a similar trajectory as in NLP. Specifically, generative pretext tasks with the masked prediction have become a de facto standard SSL practice in NLP (e.g., BERT). By contrast, early attempts at generative methods in vision have been outperformed by their discriminative counterparts (like contrastive learning). However, the success of masked image modeling has revived the autoencoder-based visual pretraining method. As a milestone to bridge the gap with BERT in NLP, masked autoencoder in vision has attracted unprecedented attention. This work conducts a survey on masked autoencoders for visual SSL.
BibTeX
@inproceedings{ijcai2023p762,
title = {A Survey on Masked Autoencoder for Visual Self-supervised Learning},
author = {Zhang, Chaoning and Zhang, Chenshuang and Song, Junha and Yi, John Seon Keun and Kweon, In So},
booktitle = {Proceedings of the Thirty-Second International Joint Conference on
Artificial Intelligence, {IJCAI-23}},
publisher = {International Joint Conferences on Artificial Intelligence Organization},
editor = {Edith Elkind},
pages = {6805--6813},
year = {2023},
month = {8},
note = {Survey Track},
doi = {10.24963/ijcai.2023/762},
url = {https://doi.org/10.24963/ijcai.2023/762},
}