2023
CAST: Cross-Attention in Space and Time for Video Action Recognition
NeurIPS 2023poster
Recognizing human actions in videos requires spatial and temporal understanding. Most existing action recognition models lack a balanced spatio-temporal understanding of videos. In this work, we propose a novel two-stream architecture, called Cross-Attention in Space and Time (CAST), that achieves a…