CM-PIE: Cross-Modal Perception for Interactive-Enhanced Audio-Visual Video Parsing
Audio-visual video parsing is the task of categorizing a video with weak labels at the segment level, and predicting them as audible or visible events. Recent methods have leveraged the attention mechanism to capture the semantic correlations among the whole video across the audio-visual modalities.…