A Separation Priority Pipeline for Single-Channel Speech Separation in Noisy Environments
Shaoxiang Dang, Tetsuya Matsumoto, Yoshinori Takeuchi, Hiroaki Kudo
Abstract
The research topic of speech separation is dedicated to addressing the issues associated with the separation of overlapping speeches, and it becomes more challenging when mixed speeches take place in noisy environments. The main pipeline so far consists of a first-place enhancement module and a following-up separation module, and some may further employ a so-called gradient modulation technique to solve the conflict that may exist during optimizing multitask of enhancement and separation. In this work, we conceive a hypothesis that separation-sensitive information might be erased during the enhancement module in traditional pipelines, and we propose a separation priority pipeline (SPP) to verify it. Furthermore, this work also provides an independent encoders and decoders scheme (IEDS) which is able to mitigate gradient conflict. According to our experiments, we found the effectiveness of SPP compared to previous work: 0.70 dB improvement of SI-SNRi, 0.74 dB improvement of SDRi, 1.52 % improvement of STOI on Libri2mix, and better generalization on LibriCSS.
BibTeX
@inproceedings{icassp2024_aseparationprior,
title = {A Separation Priority Pipeline for Single-Channel Speech Separation in Noisy Environments},
author = {Shaoxiang Dang and Tetsuya Matsumoto and Yoshinori Takeuchi and Hiroaki Kudo},
booktitle = {ICASSP 2024},
year = {2024}
}