2024
Semantic-Guided Network with Contrastive Learning for Video Caption
ICASSP 2024accepted
Video captioning is a challenging task, which aims at generating a sentence to describe the content of a video using the natural language. Many existing methods model visual features (2D/3D) extracted from videos to generate captions, but they neglect semantic guidance. Empirically, visual features…