ICASSP 2019accepted0 citations

Proximal Deep Recurrent Neural Network for Monaural Singing Voice Separation

Weitao Yuan, Shengbei Wang, Xiangrui Li, Masashi Unoki, Wenwu Wang

Abstract

The recent deep learning methods can offer state-of-the-art performance for Monaural Singing Voice Separation (MSVS). In these deep methods, the recurrent neural network (RNN) is widely employed. This work proposes a novel type of Deep RNN (DRNN), namely Proximal DRNN (P-DRNN) for MSVS, which improves the conventional Stacked RNN (S-RNN) by introducing a novel interlayer structure. The interlayer structure is derived from an optimization problem for Monaural Source Separation (MSS). Accordingly, this enables a new hierarchical processing in the proposed P-DRNN with the explicit state transfers between different layers and the skip connections from the inputs, which are efficient for source separation. Finally, the proposed approach is evaluated on the MIR-IK dataset to verify its effectiveness. The numerical results show that the P-DRNN performs better than the conventional S-RNN and several recent MSVS methods.

BibTeX
@inproceedings{icassp2019_proximaldeeprecu,
  title = {Proximal Deep Recurrent Neural Network for Monaural Singing Voice Separation},
  author = {Weitao Yuan and Shengbei Wang and Xiangrui Li and Masashi Unoki and Wenwu Wang},
  booktitle = {ICASSP 2019},
  year = {2019}
}