Complex Ratio Masking For Singing Voice Separation
Yixuan Zhang, Yuzhou Liu, DeLiang Wang
Abstract
Music source separation is important for applications such as karaoke and remixing. Much of previous research focuses on estimating short-time Fourier transform (STFT) magnitude and discarding phase information. We observe that, for singing voice separation, phase can make considerable improvement in separation quality. This paper proposes a complex ratio masking method for voice and accompaniment separation. The proposed method employs DenseUNet with self attention to estimate the real and imaginary components of STFT for each sound source. A simple ensemble technique is introduced to further improve separation performance. Evaluation results demonstrate that the proposed method outperforms recent state-of-the-art models for both separated voice and accompaniment.
BibTeX
@inproceedings{icassp2021_complexratiomask,
title = {Complex Ratio Masking For Singing Voice Separation},
author = {Yixuan Zhang and Yuzhou Liu and DeLiang Wang},
booktitle = {ICASSP 2021},
year = {2021}
}