Weakly-Supervised Video Highlight Detection by Characteristic and Commonality Modeling
Chengze Zhao, Zixuan Zhao, Xu Zhao
Abstract
Video highlight detection is important for video understanding, as it localizes the attractive regions in the video automatically. Because the fully-supervised video highlight is expensive for its frame-level annotation, we present a novel network for weakly-supervised video highlight detection based on the characteristic and commonality of key moments, utilizing the intrinsic properties of highlight moments. Specifically, the characteristic of key moments refer to the fact that the attractive segments in a video often show significant differences from the ordinary segments, i.e., non-highlight segments. However, apart from the real highlight moment showing significant difference to the ordinary, the noisy frames are also different. Therefore, commonalities among the videos within the same class is proposed to ensure the rationality of the selected moments. These proposed properties of the highlight moment are implemented as two auxiliary modules with the commonly used weakly supervised video highlight detection architecture, injecting the proposed prior information to the network. In addition, an video-level importance aggregation module is proposed to estimate the importance scores for each segment in an unified model. Our method achieves the superior performance on the commonly used benchmarks.
BibTeX
@inproceedings{icassp2025_weaklysupervised,
title = {Weakly-Supervised Video Highlight Detection by Characteristic and Commonality Modeling},
author = {Chengze Zhao and Zixuan Zhao and Xu Zhao},
booktitle = {ICASSP 2025},
year = {2025}
}