Realistic human action recognition: When deep learning meets VLAD
Lei Zhang, Yangyang Feng, Jiqing Han, Xiantong Zhen
Abstract
Human action recognition from realistic scenarios is extremely challenging due to large intra-class variation and complex background clutters. In this paper, by leveraging the strength of deep learning and vector of locally aggregated descriptors (VLAD), we propose a new methods for human action recognition from realistic datsets. We adopt stack convolu-tional independent subspace analysis (ISA) networks to learn 3D cuboid representation directly from spatio-temporal video data; we propose an improved VLAD by incorporating the spatio-temporal geometrical information to encode the deep learned local features. On two challenging realistic datasets: the YouTube action and HMDB51 datasets, the proposed method achieves state-of-the-art performance with an efficient linear SVM classifier, which is competitive with and even better than existing sophisticated algorithms.
BibTeX
@inproceedings{icassp2016_realistichumanac,
title = {Realistic human action recognition: When deep learning meets VLAD},
author = {Lei Zhang and Yangyang Feng and Jiqing Han and Xiantong Zhen},
booktitle = {ICASSP 2016},
year = {2016}
}