Linearized Kernel Representation Learning from Video Tensors by Exploiting Manifold Geometry for Gesture Recognition
Abstract
A video tensor is an organized multidimensional array of numerical values. In this paper, we explore the underlying manifold geometry of a video tensor by factorizing it using modified higher order singular value decomposition (HOSVD). Each factor (mode matrix) of a video tensor obtained after modified HOSVD can be thought of as a subspace and hence represents a point in Grassmann manifold (ℳ <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">GM</sub> ). These factors cumulatively represent a point in product Grassmann manifold (ℳ <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">PGM</sub> ). We propose a novel kernel for ℳ <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">PGM</sub> that measures the similarity between two points in ℳ <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">PGM</sub> and generates a kernel-gram matrix. For representation learning, we diagonalize the obtained kernel-gram matrix and generate a small fixed length representation corresponding to each point in ℳ <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">PGM</sub> . Classification is performed in sparse framework with minimum residual error as classifier. Experimentation is carried out over Cambridge hand gesture and UMD Keck body gesture databases for both static and dynamic settings. Experimental study shows that even with small length feature representation, there is a significant improvement in classification results as compared to state-of the-art techniques.
BibTeX
@inproceedings{icassp2019_linearizedkernel,
title = {Linearized Kernel Representation Learning from Video Tensors by Exploiting Manifold Geometry for Gesture Recognition},
author = {Krishan Sharma and Renu Rameshan},
booktitle = {ICASSP 2019},
year = {2019}
}