ICASSP 2023accepted0 citations

Modulation-Based Center Alignment and Motion Mining for Spatial Temporal Action Detection

Weiji Zhao, Kefeng Huang, Chongyang Zhang

Abstract

The goal of spatial-temporal action detection is to generate spatial-temporally aligned action tubes. Most of the existing 2D CNN-based solutions directly aggregate temporal adjacent contexts through frames without alignment. The misaligned spatial-temporal contextual features might lead to chaotic representation and misaligned action tubes. Moreover, most existing methods fail to efficiently exploit motion dependencies. In this paper, we propose Modulation-based Center Alignment (MCA) and Sparse Valuable Motion Mining (SVMM) for more accurate action detection: With deformable convolution, key-frame based modulation is firstly designed to align the action center between temporal frames; then motion region guided sparse self-attention is developed for valuable motion mining. Our framework can outperform current 2D CNN-based methods significantly, based on the experimental result on two widely used benchmarks of JH-MDB and UCF101-24.

BibTeX
@inproceedings{icassp2023_modulationbasedc,
  title = {Modulation-Based Center Alignment and Motion Mining for Spatial Temporal Action Detection},
  author = {Weiji Zhao and Kefeng Huang and Chongyang Zhang},
  booktitle = {ICASSP 2023},
  year = {2023}
}
Modulation-Based Center Alignment and Motion Mining for Spatial Temporal Action Detection · ICASSP 2023