EMNLP 2023long findings0 citations

Sparse Frame Grouping Network with Action Centered for Untrimmed Video Paragraph Captioning

Guorui Yu, Yimin Hu, Yuejie Zhang, Rui Feng, Tao Zhang, Shang Gao

Abstract

Generating paragraph captions for untrimmed videos without event annotations is challenging, especially when aiming to enhance precision and minimize repetition at the same time. To address this challenge, we propose a module called Sparse Frame Grouping (SFG). It dynamically groups event information with the help of action information for the entire video and excludes redundant frames within pre-defined clips. To enhance the performance, an Intra Contrastive Learning technique is designed to align the SFG module with the core event content in the paragraph, and an Inter Contrastive Learning technique is employed to learn action-guided context with reduced static noise simultaneously. Extensive experiments are conducted on two benchmark datasets (ActivityNet Captions and YouCook2). Results demonstrate that SFG outperforms the state-of-the-art methods on all metrics.

video paragraph captioningtransformergroupingaction centeredcontrastive learning
BibTeX
@inproceedings{
yu2023sparse,
title={Sparse Frame Grouping Network with Action Centered for Untrimmed Video Paragraph Captioning},
author={Guorui Yu and Yimin Hu and Yuejie Zhang and Rui Feng and Tao Zhang and Shang Gao},
booktitle={The 2023 Conference on Empirical Methods in Natural Language Processing},
year={2023},
url={https://openreview.net/forum?id=5K2fiOlcGG}
}
Sparse Frame Grouping Network with Action Centered for Untrimmed Video Paragraph Captioning · EMNLP 2023