2026
How Far Can We Go With Synthetic Data for Audio-Visual Sound Source Localization?
CVPR 2026
We present the first scalable framework for training sound source localization (SSL) models using synthetic data from text-to-X models. Although SSL has made notable progress, existing models remain constrained by limited-scale, uncurated real-world datasets that often suffer from semantic misalignm