On SDW-MWF and Variable Span Linear Filter with Application to Speech Recognition in Noisy Environments
Ziteng Wang, Lu Yin, Junfeng Li, Yonghong Yan
Abstract
Neural network based spectral mask estimation for acoustic beamforming, which consists of linear filtering and mask estimation, has shown to be a promising approach for robust speech recognition in noisy environments. Nevertheless, few improvements are made on the linear filtering. In this paper, we investigate the Speech Distortion Weighted Multichannel Wiener Filter (SDW-MWF) and the variable span linear filter, and prove that they can be linked by Generalized Eigenvalue Decomposition (GEVD) of the speech covariance matrix. The resulting GEVD based SDW-MWF largely reduces the word error rate and even achieves competitive recognition performance with the state-of-the-art generalized eigenvalue beamformer. Furthermore, we found that the recent signal approximation is no better than mask approximation when combined in calculating the linear filter coefficients.
BibTeX
@inproceedings{icassp2018_onsdwmwfandvaria,
title = {On SDW-MWF and Variable Span Linear Filter with Application to Speech Recognition in Noisy Environments},
author = {Ziteng Wang and Lu Yin and Junfeng Li and Yonghong Yan},
booktitle = {ICASSP 2018},
year = {2018}
}