← Search

Sung Jin Um

3 accepted papers

2025

Object-aware Sound Source Localization via Audio-Visual Scene Understanding

CVPR 2025poster

Audio-visual sound source localization task aims to spatially localize sound-making objects within visual scenes by integrating visual and audio cues. However, existing methods struggle with accurately localizing sound-making objects in complex scenes, particularly when visually similar silent objec…

2025

Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection

AAAI 2025technical

The goal of video moment retrieval and highlight detection is to identify specific segments and highlights based on a given text query. With the rapid growth of video content and the overlap between these tasks, recent works have addressed both simultaneously. However, they still struggle to fully c…

2024

Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge

CVPR 2024poster

The goal of the multi-sound source localization task is to localize sound sources from the mixture individually. While recent multi-sound source localization methods have shown improved performance they face challenges due to their reliance on prior information about the number of objects to be sepa…