ICLR 2026poster0 citations

Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source Localization

Min-Sang Baek, Gyeong-Su Kim, Donghyun Kim, Joon-Hyuk Chang

Abstract

Sound source localization (SSL) is a fundamental task in spatial audio understanding, yet most deep neural network-based methods are constrained by fixed array geometries and predefined directional grids, limiting generalizability and scalability. To address these issues, we propose _audio-geometry-grid representation learning_ (AGG-RL), a novel framework that jointly learns audio-geometry and grid representations in a shared latent space, enabling both geometry-invariant and grid-flexible SSL. Moreover, to enhance generalizability and interpretability, we introduce two physics-informed components: a _learnable non-uniform discrete Fourier transform_ (LNuDFT), which optimizes the dense allocation of frequency bins in a non-uniform manner to emphasize informative phase regions, and a _relative microphone positional encoding_ (rMPE), which encodes relative microphone coordinates in accordance with the nature of inter-channel time differences. Experiments on synthetic and real datasets demonstrate that AGG-RL achieves superior performance, particularly under unseen conditions. The results highlight the potential of representation learning with physics-informed design towards a universal solution for spatial acoustic scene understanding across diverse scenarios.

Sound Source LocalizationGeometry-InvariantGrid-FlexibleRepresentation LearningPhysics-Informed DesignLearnable Non-uniform DFTRelative Microphone Positional Encoding
BibTeX
@inproceedings{
baek2026physicsinformed,
title={Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source Localization},
author={Min-Sang Baek and Gyeong-Su Kim and Donghyun Kim and Joon-Hyuk Chang},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=bWXpJFesLS}
}
Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source Localization · ICLR 2026