← Search

Taekyung Kim*

2 accepted papers

2024

Learning with Unmasked Tokens Drives Stronger Vision Learners

ECCV 2024poster

"Masked image modeling (MIM) has become a leading self-supervised learning strategy. MIMs such as Masked Autoencoder (MAE) learn strong representations by randomly masking input tokens for the encoder to process, with the decoder reconstructing the masked tokens to the input. However, MIM pre-traine…

2024

Leveraging temporal contextualization for video action recognition

ECCV 2024poster

"We propose a novel framework for video understanding, called (), which leverages essential temporal information through global interactions in a spatio-temporal domain within a video. To be specific, we introduce Temporal Contextualization (TC), a layer-wise temporal information infusion mechanism…