NeurIPS 2019poster127 citations

LiteEval: A Coarse-to-Fine Framework for Resource Efficient Video Recognition

Zuxuan Wu, Caiming Xiong, Yu-Gang Jiang, Larry S. Davis

Abstract

This paper presents LiteEval, a simple yet effective coarse-to-fine framework for resource efficient video recognition, suitable for both online and offline scenarios. Exploiting decent yet computationally efficient features derived at a coarse scale with a lightweight CNN model, LiteEval dynamically decides on-the-fly whether to compute more powerful features for incoming video frames at a finer scale to obtain more details. This is achieved by a coarse LSTM and a fine LSTM operating cooperatively, as well as a conditional gating module to learn when to allocate more computation. Extensive experiments are conducted on two large-scale video benchmarks, FCVID and ActivityNet, and the results demonstrate LiteEval requires substantially less computation while offering excellent classification accuracy for both online and offline predictions.

BibTeX
@inproceedings{NEURIPS2019_bd853b47,
 author = {Wu, Zuxuan and Xiong, Caiming and Jiang, Yu-Gang and Davis, Larry S},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {H. Wallach and H. Larochelle and A. Beygelzimer and F. d\textquotesingle Alch\'{e}-Buc and E. Fox and R. Garnett},
 pages = {},
 publisher = {Curran Associates, Inc.},
 title = {LiteEval: A Coarse-to-Fine Framework for Resource Efficient Video Recognition},
 url = {https://proceedings.neurips.cc/paper_files/paper/2019/file/bd853b475d59821e100d3d24303d7747-Paper.pdf},
 volume = {32},
 year = {2019}
}
LiteEval: A Coarse-to-Fine Framework for Resource Efficient Video Recognition · NeurIPS 2019