← Search

Di Hu*

4 accepted papers

2024

Can Textual Semantics Mitigate Sounding Object Segmentation Preference?

ECCV 2024poster

"The Audio-Visual Segmentation (AVS) task aims to segment sounding objects in the visual space using audio cues. However, in this work, it is recognized that previous AVS methods show a heavy reliance on detrimental segmentation preferences related to audible objects, rather than precise audio guida…

2024

Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes

ECCV 2024poster

"Traditional reference segmentation tasks have predominantly focused on silent visual scenes, neglecting the integral role of multimodal perception and interaction in human experiences. In this work, we introduce a novel task called Reference Audio-Visual Segmentation (Ref-AVS), which seeks to segme…

2024

Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation

ECCV 2024poster

"Audio-Visual Segmentation (AVS) aims to achieve pixel-level localization of sound sources in videos, while Audio-Visual Semantic Segmentation (AVSS), as an extension of AVS, further pursues semantic understanding of audio-visual scenes. However, since the AVSS task requires the establishment of aud…