End-to-End Sound Source Enhancement Using Deep Neural Network in the Modified Discrete Cosine Transform Domain
Yuma Koizumi, Noboru Harada, Yoichi Haneda, Yusuke Hioka, Kazunori Kobayashi
Abstract
This paper presents an end-to-end deep neural network (DNN)-based source enhancement on the basis of a time-frequency (T-F) mask processing in the modified discrete cosine transform (MDCT)-domain. To retrieve the target signal perfectly in the discrete Fourier transform (DFT)-domain, both amplitude and phase of the spectrum need to be manipulated. However, since it is difficult to deal with complex values by neural network straightforward way, a real-valued T-F mask is commonly estimated and only amplitude spectrum is manipulated. In this study, we use the MDCT instead of the DFT and estimate real-valued T-F masks in the MDCT-domain. The perfect retrieval can be achieved by manipulating only the real-valued MDCT-spectra. To reduce time-domain aliasing arises from manipulating the MDCT spectrum, we build an end-to-end DNN-based source enhancement using T-F mask and train the DNN to minimize an objective function defined in the time-domain. In experiments using several kinds of objective sound quality scores, we observed that the scores were significantly improved.
BibTeX
@inproceedings{icassp2018_endtoendsoundsou,
title = {End-to-End Sound Source Enhancement Using Deep Neural Network in the Modified Discrete Cosine Transform Domain},
author = {Yuma Koizumi and Noboru Harada and Yoichi Haneda and Yusuke Hioka and Kazunori Kobayashi},
booktitle = {ICASSP 2018},
year = {2018}
}