← Search

Heeseung Kwon

4 accepted papers

2021

Learning Self-Similarity in Space and Time As Generalized Motion for Video Action Recognition

ICCV 2021poster

Spatio-temporal convolution often fails to learn motion dynamics in videos and thus an effective motion representation is required for video understanding in the wild. In this paper, we propose a rich and robust motion representation based on spatio-temporal self-similarity (STSS). Given a sequence…

Cited by 56PDFcodeScholar
2021

Relational Self-Attention: What's Missing in Attention for Video Understanding

NeurIPS 2021poster

Convolution has been arguably the most important feature transform for modern neural networks, leading to the advance of deep learning. Recent emergence of Transformer networks, which replace convolution layers with self-attention blocks, has revealed the limitation of stationary convolution kerne…

2020

MotionSqueeze: Neural Motion Feature Learning for Video Understanding

ECCV 2020poster

Motion plays a crucial role in understanding videos and most state-of-the-art neural models for video classification incorporate motion information typically using optical flows extracted by a separate off-the-shelf method. As the frame-by-frame optical flows require heavy computation, incorporating…

Cited by 174SourcePDFScholar