Large Scale Holistic Video Understanding
Ali Diba, Mohsen Fayyaz, Vivek Sharma, Manohar Paluri, Jürgen Gall, Rainer Stiefelhagen, Luc Van Gool
Abstract
Video recognition has been advanced in recent years by benchmarks with rich annotations. However, research is still mainly limited to human action or sports recognition - focusing on a highly specific video understanding task and thus leaving a significant gap towards describing the overall content of a video. We fill this gap by presenting a large-scale ``Holistic Video Understanding Dataset""~(HVU). HVU is organized hierarchically in a semantic taxonomy that focuses on multi-label and multi-task video understanding as a comprehensive problem that encompasses the recognition of multiple semantic aspects in the dynamic scene. HVU contains approx.~572k videos in total with 9 million annotations for training, validation, and test set spanning over 3142 labels. HVU encompasses semantic aspects defined on categories of scenes, objects, actions, events, attributes, and concepts which naturally capture the real-world scenarios.
BibTeX
@inproceedings{eccv2020_largescaleholist,
title = {Large Scale Holistic Video Understanding},
author = {Ali Diba and Mohsen Fayyaz and Vivek Sharma and Manohar Paluri and Jürgen Gall and Rainer Stiefelhagen and Luc Van Gool},
booktitle = {ECCV 2020},
year = {2020}
}