ICASSP 2022accepted0 citations

An End-to-End Deep Learning Framework For Multiple Audio Source Separation And Localization

Yu Chen, Bowen Liu, Zijian Zhang, Hun-Seok Kim

Abstract

Sound source separation and localization for situational awareness enables a wide range of applications such as hearing enhancement and audio beam-forming. We present an end-to-end deep learning framework to separate and localize multiple audio sources from the mixture of multi-channels. The proposed framework jointly estimates the separated sources and their time difference of arrival (TDOA) at different microphones, then it obtains the direction-of-arrival (DOA) for each source. A new structure to reconstruct the mixed signal is introduced for joint optimization of source separation and TDOA estimation. In addition, a discriminator network is added during the training phase to further improve the separation quality. Experiment results demonstrate that the proposed method achieves state-of-the-art accuracy on source separation as well as DOA estimation.

BibTeX
@inproceedings{icassp2022_anendtoenddeeple,
  title = {An End-to-End Deep Learning Framework For Multiple Audio Source Separation And Localization},
  author = {Yu Chen and Bowen Liu and Zijian Zhang and Hun-Seok Kim},
  booktitle = {ICASSP 2022},
  year = {2022}
}