2024
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
ICML 2024poster
Masked Video Autoencoder (MVA) approaches have demonstrated their potential by significantly outperforming previous video representation learning methods. However, they waste an excessive amount of computations and memory in predicting uninformative tokens/frames due to random masking strategies. (e…