M2SUM: Multi-Granularity Scale-Adaptive Video Summarizer towards Informative Context Representation Learning
Yunzuo Zhang, Yameng Liu, Weili Kang
Abstract
Video summarization intends to automatically select meaningful segments from untrimmed videos. Although previous efforts have achieved remarkable progress, they still struggle to robustly aggregate and effectively process multi-granularity contextual information within videos, which hinders understanding towards video content. To address these issues, we propose M2SUM, which is composed of three dominant components including the embedding learning attention (ELA) module, multi-granularity aggregator (MGA), and semantic scale-adaption (SSA) module. ELA dynamically enhances pre-trained visual features by considering the similarity relationship across frame-level and video-level embeddings. MGA incorporates self-attention and temporal convolution into a unified learnable module, robustly learning long-range and short-range multi-granularity temporal dependencies. SSA is exploited to adaptively perform representation fusion after deep interaction across multi-granularity temporal dependencies. According to the fused representations, M2SUM predicts importance scores and generates video summaries. Extensive experiments on standard datasets have proved the effectiveness and superiority of our method in F-score and rank-based evaluations.
BibTeX
@inproceedings{icassp2024_m2summultigranul,
title = {M2SUM: Multi-Granularity Scale-Adaptive Video Summarizer towards Informative Context Representation Learning},
author = {Yunzuo Zhang and Yameng Liu and Weili Kang},
booktitle = {ICASSP 2024},
year = {2024}
}