← Search

Hengyi Hong

3 accepted papers

2026

Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning

ICASSP 2026oral

We present Task 5 of the DCASE 2025 Challenge: an Audio Question Answering (AQA) benchmark spanning multiple domains of sound understanding. This task defines three QA subsets (Bioacoustics, Temporal Soundscapes, and Complex QA) to test audio-language models on interactive question-answering over di…

Cited by 0SourcePDFScholar
2025

An Experimental Study on Joint Modeling for Sound Event Localization and Detection with Source Distance Estimation

ICASSP 2025accepted

In traditional sound event localization and detection (SELD) tasks, the focus is typically on sound event detection (SED) and direction-of-arrival (DOA) estimation, but they fall short of providing full spatial information about the sound source. The 3D SELD task addresses this limitation by integra…

Cited by 0SourceScholar
2025

MVANet: Multi-Stage Video Attention Network for Sound Event Localization and Detection with Source Distance Estimation

ICASSP 2025accepted

Sound event localization and detection with source distance estimation (3D SELD) involves not only identifying the sound category and its direction-of-arrival (DOA) but also predicting the source's distance, aiming to provide full information about the sound position. This paper proposes a multi-sta…

Cited by 0SourceScholar