← Search

Dan Guo*

2 accepted papers

2024

Label-anticipated Event Disentanglement for Audio-Visual Video Parsing

ECCV 2024poster

"Audio-Visual Video Parsing (AVVP) task aims to detect and temporally locate events within audio and visual modalities. Multiple events can overlap in the timeline, making identification challenging. While traditional methods usually focus on improving the early audio-visual encoders to embed more e…

Cited by 15SourcePDFScholar
2024

Training A Small Emotional Vision Language Model for Visual Art Comprehension

ECCV 2024poster

"This paper develops small vision language models to understand visual art, which, given an art work, aims to identify its emotion category and explain this prediction with natural language. While small models are computationally efficient, their capacity is much limited compared with large models.…