Spatial Annotation-free Training for Sound Event Localization and Detection
Masahiro Yasuda, Shoichiro Saito, Nao Sato, Noboru Harada
Abstract
Sound Event Localization and Detection (SELD) is the task of estimating the class, duration, and direction of arrival (DOA) of sound events. State-of-the-art SELD systems use a data-driven approach based on Deep Neural Networks (DNNs) to deal with complex situations involving overlapping and moving sound sources. Such systems need to be trained on real-recording data to succeed in real-world situations. However, annotation of real-world data, especially spatial annotation, incurs huge costs, and the amount of labeled data is currently insufficiently small. Therefore, we are introducing spatial annotation-free training for SELD, which trains a SELD system using only sound class and duration labels. As a first attempt at this task, we propose Beam-based Multiple Instance Learning (Beam-MIL). Beam-MIL first instantiates acoustic signals for each DOA by beamforming. Then, the sound events contained in each instance are trained indirectly by MIL without using each instance’s ground truth information, i.e., spatial annotation. Experimental results show that Beam-MIL can effectively train a valid DOA estimator without spatial annotations. Moreover, when the small amount of the annotated data is available, enlarging data size by adding annotation-free data significantly improved the performance of the system.
BibTeX
@inproceedings{icassp2025_spatialannotatio,
title = {Spatial Annotation-free Training for Sound Event Localization and Detection},
author = {Masahiro Yasuda and Shoichiro Saito and Nao Sato and Noboru Harada},
booktitle = {ICASSP 2025},
year = {2025}
}