DNN-based speech mask estimation for eigenvector beamforming
Lukas Pfeifenberger, Matthias Zöhrer, Franz Pernkopf
Abstract
In this paper, we present an optimal multi-channel Wiener filter, which consists of an eigenvector beamformer and a single-channel postfilter. We show that both components solely depend on a speech presence probability, which we learn using a deep neural network, consisting of a deep autoencoder and a softmax regression layer. To prevent the DNN from learning specific speaker and noise types, we do not use the signal energy as input feature, but rather the cosine distance between the dominant eigenvectors of consecutive frames of the power spectral density of the noisy speech signal. We compare our system against the BeamformIt toolkit, and state-of-the-art approaches such as the front-end of the best system of the CHiME3 challenge. We show that our system yields superior results, both in terms of perceptual speech quality and classification error.
BibTeX
@inproceedings{icassp2017_dnnbasedspeechma,
title = {DNN-based speech mask estimation for eigenvector beamforming},
author = {Lukas Pfeifenberger and Matthias Zöhrer and Franz Pernkopf},
booktitle = {ICASSP 2017},
year = {2017}
}