Text-Infused Audio-Visual Video Parsing with Semantic-Aware Multimodal Contrastive Learning
Pengcheng Zhao, Yanxiang Chen, Dan Guo, Yuanzhi Yao
Abstract
The Audio-Visual Video Parsing task aims to recognize events occurring in video segments for each modality. Presently, the excellent performance in handling video parsing is shown by generating pseudo labels at the segment level. However, these approaches still suffer from adequate semantic learning of fine-grained segment features, which can cause errors in the prediction of those ambiguous events and affect the predictions of another modality. To tackle this issue, we propose a novel Text-Infused Parsing Network (TIPNet). Specifically, the event text modality, which can offer more precise semantics than the audio and visual modalities, is introduced to enhance the event-related audio/visual segment features by encoding the cross-modal interactions. Furthermore, to obtain more discriminative segment features for each modality, we propose a novel multimodal contrastive loss with a semantic-aware weighting mechanism. Experimental results on the benchmark dataset demonstrate the superior performance of our approach compared to the state-of-the-art methods.
BibTeX
@inproceedings{icassp2025_textinfusedaudio,
title = {Text-Infused Audio-Visual Video Parsing with Semantic-Aware Multimodal Contrastive Learning},
author = {Pengcheng Zhao and Yanxiang Chen and Dan Guo and Yuanzhi Yao},
booktitle = {ICASSP 2025},
year = {2025}
}