2024
Weakly-Supervised Audio-Visual Video Parsing with Prototype-based Pseudo-Labeling
CVPR 2024poster
In this paper we address the weakly-supervised Audio-Visual Video Parsing (AVVP) problem which aims at labeling events in a video as audible visible or both and temporally localizing and classifying them into known categories. This is challenging since we only have access to video-level (weak) event…