Robust and compact video descriptor learned by deep neural network
Abstract
In this paper, we propose to extract robust video descriptor by training deep neural network to automatically capture the intrinsic visual characteristics of digital video. More specifically, we first train a conditional generative model to capture the spatio-temporal correlations among visual contents and represent them as an intermediate descriptor. A nonlinear encoder, with the functions of dimension reduction and error correcting, is then trained to learn a compressed yet more robust representation of the intermediate descriptor. The cascade of the conditional generative model and the encoder constitutes the building block of the deep network for learning video descriptor. As a post-processing component, the top layers of the network are trained to optimize the robustness and discriminative capability of the output descriptor. Experimental results on benchmark databases confirm that the descriptor learned by deep neural network shows excellent robustness against photometric, geometric, temporal and combined distortions, and it can attain an F <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</inf> score of 0.982 in content identification, which is much higher than hand-engineered descriptors.
BibTeX
@inproceedings{icassp2017_robustandcompact,
title = {Robust and compact video descriptor learned by deep neural network},
author = {Yue-Nan Li and Xue Piao Chen},
booktitle = {ICASSP 2017},
year = {2017}
}