ICASSP 2018accepted0 citations

Perceptually Guided Speech Enhancement Using Deep Neural Networks

Yan Zhao, Buye Xu, Ritwik Giri, Tao Zhang

Abstract

Human listeners often have difficulties understanding speech in the presence of background noise in the real world. Recently, supervised learning based speech enhancement approaches have achieved substantial success, and show significant improvements over the conventional approaches. However, existing supervised learning based approaches often try to minimize the mean squared error between the enhanced output and the pre-defined training target (e.g., the log power spectrum of clean speech), even though the purpose of such speech enhancement is to improve speech understanding in noise. In this paper, we propose a new deep neural networks based enhancement approach by incorporating a speech perception model into the loss function. Specifically, we use the short-time objective intelligibility metric in the loss in addition to the mean squared error. Optimizing the proposed perceptually guided loss is expected to improve speech intelligibility further. Systematic evaluations show that our proposed approach is able to improve speech intelligibility in a wide range of signal-to-noise ratios and noise types while maintaining speech quality.

BibTeX
@inproceedings{icassp2018_perceptuallyguid,
  title = {Perceptually Guided Speech Enhancement Using Deep Neural Networks},
  author = {Yan Zhao and Buye Xu and Ritwik Giri and Tao Zhang},
  booktitle = {ICASSP 2018},
  year = {2018}
}