← Search

Jinxing Zhou*

1 accepted papers

2024

Label-anticipated Event Disentanglement for Audio-Visual Video Parsing

ECCV 2024poster

"Audio-Visual Video Parsing (AVVP) task aims to detect and temporally locate events within audio and visual modalities. Multiple events can overlap in the timeline, making identification challenging. While traditional methods usually focus on improving the early audio-visual encoders to embed more e…

Cited by 15SourcePDFScholar