ICASSP 2018accepted0 citations

Mask Weighted Stft Ratios for Relative Transfer Function Estimation and ITS Application to Robust ASR

Zhong-Qiu Wang, DeLiang Wang

Abstract

Deep learning based single-channel time-frequency (T-F) masking has shown considerable potential for beamforming and robust ASR. This paper proposes a simple but novel relative transfer function (RTF) estimation algorithm for microphone arrays, where the RTF between a reference signal and a non-reference signal at each frequency band is estimated as a weighted average of the ratios of the two STFT (short-time Fourier transform) coefficients of the speech-dominant T-F units. Similarly, the noise covariance matrix is estimated from noise-dominant T-F units. An MVDR beamformer is then constructed for robust ASR. Experiments on the two- and six-channel track of the CHiME-4 challenge show consistent improvement over a weighted delay-and-sum (WDAS) beamformer, a generalized eigenvector beamformer, a parameterized multi-channel Wiener filter, an MVDR beamformer based on conventional direction of arrival (DOA) estimation, and two MVDR beamformers both based on eigendecomposition.

BibTeX
@inproceedings{icassp2018_maskweightedstft,
  title = {Mask Weighted Stft Ratios for Relative Transfer Function Estimation and ITS Application to Robust ASR},
  author = {Zhong-Qiu Wang and DeLiang Wang},
  booktitle = {ICASSP 2018},
  year = {2018}
}
Mask Weighted Stft Ratios for Relative Transfer Function Estimation and ITS Application to Robust ASR · ICASSP 2018