ICASSP 2024accepted0 citations

A Separation Priority Pipeline for Single-Channel Speech Separation in Noisy Environments

Shaoxiang Dang, Tetsuya Matsumoto, Yoshinori Takeuchi, Hiroaki Kudo

Abstract

The research topic of speech separation is dedicated to addressing the issues associated with the separation of overlapping speeches, and it becomes more challenging when mixed speeches take place in noisy environments. The main pipeline so far consists of a first-place enhancement module and a following-up separation module, and some may further employ a so-called gradient modulation technique to solve the conflict that may exist during optimizing multitask of enhancement and separation. In this work, we conceive a hypothesis that separation-sensitive information might be erased during the enhancement module in traditional pipelines, and we propose a separation priority pipeline (SPP) to verify it. Furthermore, this work also provides an independent encoders and decoders scheme (IEDS) which is able to mitigate gradient conflict. According to our experiments, we found the effectiveness of SPP compared to previous work: 0.70 dB improvement of SI-SNRi, 0.74 dB improvement of SDRi, 1.52 % improvement of STOI on Libri2mix, and better generalization on LibriCSS.

BibTeX
@inproceedings{icassp2024_aseparationprior,
  title = {A Separation Priority Pipeline for Single-Channel Speech Separation in Noisy Environments},
  author = {Shaoxiang Dang and Tetsuya Matsumoto and Yoshinori Takeuchi and Hiroaki Kudo},
  booktitle = {ICASSP 2024},
  year = {2024}
}
A Separation Priority Pipeline for Single-Channel Speech Separation in Noisy Environments · ICASSP 2024