← Search

Hyeonggon Ryu

9 accepted papers

2026

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

AAAI 2026technical

Audio comprehension—including speech, non-speech sounds, and music—is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challen

Cited by 0SourcePDFScholar
2026

Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions

CVPR 2026

We address the problem of tactile localization, where the goal is to identify image regions that share the same material properties as a tactile input. Existing visuo-tactile methods rely on global alignment and thus fail to capture the fine-grained local correspondences required for this task. The

Cited by 0SourcecodeScholar
2025

Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual Scenes

CVPR 2025poster

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited to handling either speech or non-speech sounds independently…

Cited by 0SourcePDFScholar
2024

Speech Guided Masked Image Modeling for Visually Grounded Speech

ICASSP 2024accepted

The objective of this study is to investigate the learning process of Visually Grounded Speech (VGS) models through joint learning that combines contrastive learning and masked image modeling. Typically, VGS models ahn to establish audio-visual alignment between images and then spoken captions withi…

Cited by 0SourceScholar
2023

Generative Bias for Robust Visual Question Answering

CVPR 2023poster

The task of Visual Question Answering (VQA) is known to be plagued by the issue of VQA models exploiting biases within the dataset to make its final prediction. Various previous ensemble based debiasing methods have been proposed where an additional model is purposefully trained to be biased in orde…

2023

Hindi as a Second Language: Improving Visually Grounded Speech with Semantically Similar Samples

ICASSP 2023accepted

The objective of this work is to explore the learning of visually grounded speech models (VGS) from multilingual perspective. Bilingual VGS models are generally trained with an equal number of spoken captions from both languages. However, in reality, there can be an imbalance among the languages for…

Cited by 0SourceScholar
2023

Sound Source Localization is All about Cross-Modal Alignment

ICCV 2023poster

Humans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainly explored the problem from a localization perspective. However, prior arts and existing benchmarks do not account for…

Cited by 18PDFScholar
2022

Learning Sound Localization Better from Semantically Similar Samples

ICASSP 2022accepted

The objective of this work is to localize the sound sources in visual scenes. Existing audio-visual works employ contrastive learning by assigning corresponding audio-visual pairs from the same source as positives while randomly mismatched pairs as negatives. However, these negative pairs may contai…

Cited by 0SourceScholar