Realistic human action recognition: When CNNS meet LDS
Lei Zhang, Yangyang Feng, Xuezhi Xiang, Xiantong Zhen
Abstract
In this paper, we proposed new framework for human action representation, which leverages the strengths of convolutional neural networks (CNNs) and the linear dynamical system (LDS) to represent both spatial and temporal structures of actions in videos. We make two principal contributions: first, we incorporate image-trained CNNs to detect action clip concepts, which takes advantage of different levels of information by combining the two layers in CNNs trained from images; Second, we further propose adopting a linear dynamical system (LDS) to model the relationships between these clip concepts, which captures temporal structures of actions. We have applied the proposed method on two challenging realistic benchmark datasets, and our method achieves high performance up to 86.16% on the YouTube and 82.76% UCF50 datasets, which largely outperforms most of the state-of-the-art algorithms with more sophisticated techniques.
BibTeX
@inproceedings{icassp2017_realistichumanac,
title = {Realistic human action recognition: When CNNS meet LDS},
author = {Lei Zhang and Yangyang Feng and Xuezhi Xiang and Xiantong Zhen},
booktitle = {ICASSP 2017},
year = {2017}
}