A new uncertainty decoding scheme for DNN-HMM hybrid systems with multichannel speech enhancement
Christian Huemmer, Andreas Schwarz, Roland Maas, Hendrik Barfuss, Ramón Fernandez Astudillo, Walter Kellermann
Abstract
Uncertainty decoding combines a probabilistic feature description with the acoustic model of a speech recognition system. For DNN-HMM hybrid systems, this can be realized by averaging the DNN outputs produced by a finite set of feature samples (drawn from an estimated probability distribution). In this article, we employ this sampling approach in combination with a multi-microphone speech enhancement system. We propose a new strategy for generating feature samples from multichannel signals, based on modeling the spatial coherence estimates between different microphone pairs as realizations of a latent random variable. From each coherence estimate, a spectral enhancement gain is computed and an enhanced feature vector is obtained, thus producing a finite set of feature samples, of which we average the respective DNN outputs. In the experimental part, this new uncertainty decoding strategy is shown to consistently improve the recognition accuracy of a DNN-HMM hybrid system for the 8-channel REVERB Challenge task.
BibTeX
@inproceedings{icassp2016_anewuncertaintyd,
title = {A new uncertainty decoding scheme for DNN-HMM hybrid systems with multichannel speech enhancement},
author = {Christian Huemmer and Andreas Schwarz and Roland Maas and Hendrik Barfuss and Ramón Fernandez Astudillo and Walter Kellermann},
booktitle = {ICASSP 2016},
year = {2016}
}