ICASSP 2023accepted0 citations

Efficient Compressed Video Action Recognition Via Late Fusion with a Single Network

Hayato Terao, Wataru Noguchi, Hiroyuki Iizuka, Masahito Yamamoto

Abstract

Compressed video action recognition is an action recognition approach that can achieve efficient inference by directly classifying video data obtained from multiple features stored in compressed videos. Most conventional methods use multiple networks to process compressed video features. They explore the use of lightweight networks without affecting the classification performance to reduce the computational complexity of compressed video action recognition. This study explores another approach to reduce the computational complexity, by using a single network instead of multiple networks to process compressed video features. Although training a single network cannot yield the practical classification performance of conventional methods, we propose an extended MIMO training method for action recognition to simultaneously process different features. Our experiments demonstrate that our method is the most efficient in terms of computational complexity and can achieve classification performances comparable to conventional methods.

BibTeX
@inproceedings{icassp2023_efficientcompres,
  title = {Efficient Compressed Video Action Recognition Via Late Fusion with a Single Network},
  author = {Hayato Terao and Wataru Noguchi and Hiroyuki Iizuka and Masahito Yamamoto},
  booktitle = {ICASSP 2023},
  year = {2023}
}
Efficient Compressed Video Action Recognition Via Late Fusion with a Single Network · ICASSP 2023