ICASSP 2016accepted0 citations

Combining non-negative matrix factorization and deep neural networks for speech enhancement and automatic speech recognition

Thanh T. Vu, Benjamin Bigot, Eng Siong Chng

Abstract

Sparse Non-negative Matrix Factorization (SNMF) and Deep Neural Networks (DNN) have emerged individually as two efficient machine learning techniques for single-channel speech enhancement. Nevertheless, there are only few works investigating the combination of SNMF and DNN for speech enhancement and robust Automatic Speech Recognition (ASR). In this paper, we present a novel combination of speech enhancement components based-on SNMF and DNN into a full-stack system. We refine the cost function of the DNN to back-propagate the reconstruction error of the enhanced speech. Our proposal is compared with several state-of-the-art speech enhancement systems. Evaluations are conducted on the data of CHiME-3 challenge which consists of real noisy speech recordings captured under challenging noisy conditions. Our system yields significant improvements for both objective quality speech enhancement measurements with relative gain of 30%, and a 10% relative Word Error Rate reduction for ASR compared to the best baselines.

BibTeX
@inproceedings{icassp2016_combiningnonnega,
  title = {Combining non-negative matrix factorization and deep neural networks for speech enhancement and automatic speech recognition},
  author = {Thanh T. Vu and Benjamin Bigot and Eng Siong Chng},
  booktitle = {ICASSP 2016},
  year = {2016}
}
Combining non-negative matrix factorization and deep neural networks for speech enhancement and automatic speech recognition · ICASSP 2016