2020
RubiksNet: Learnable 3D-Shift for Efficient Video Action Recognition
ECCV 2020poster
Video action recognition is a complex task dependent on modeling spatial and temporal context. Standard approaches rely on 2D or 3D convolutions to process such context, resulting in expensive operations with millions of parameters. Recent efficient architectures leverage a channel-wise shift-based…