Hierarchical Spatial-Temporal Transformer with Motion Trajectory for Individual Action and Group Activity Recognition
Xiaolin Zhu, Dongli Wang, Yan Zhou
Abstract
Group activity recognition, which aims to simultaneously understand individual action and group activity in video clips, plays a fundamental role in video analysis. In this paper, we propose a novel reasoning network, Hierarchical Spatial-Temporal Transformer termed HSTT, for individual action and group activity recognition, which focuses on capturing the various degrees of spatial-temporal dynamic interactions adaptively and jointly among actors. Specifically, we first design a hierarchical spatial-temporal Transformer by capturing different levels of relationships to deal with unequal interaction relationships among actors. Furthermore, our proposed spatial-temporal Transformer (STT) block is capable of fully mining long-range spatial-temporal interactions with the virtue of the merge function and cross attention mechanism. Besides, we adopt the motion trajectory branch to provide complementary dynamic features for improving recognition performance. Extensive experiments on the two public GAR datasets clearly show that our approach can achieve very competitive performance by comparing them with state-of-the-art works.
BibTeX
@inproceedings{icassp2023_hierarchicalspat,
title = {Hierarchical Spatial-Temporal Transformer with Motion Trajectory for Individual Action and Group Activity Recognition},
author = {Xiaolin Zhu and Dongli Wang and Yan Zhou},
booktitle = {ICASSP 2023},
year = {2023}
}