ICASSP 2019accepted0 citations

Joint Training of Complex Ratio Mask Based Beamformer and Acoustic Model for Noise Robust Asr

Yong Xu, Chao Weng, Like Hui, Jianming Liu, Meng Yu, Dan Su, Dong Yu

Abstract

In this paper, we present a joint training framework between the multi-channel beamformer and the acoustic model for noise robust automatic speech recognition (ASR). The complex ratio mask (CRM), demonstrated to be more effective than the ideal ratio mask (IRM), is proposed to estimate the covariance matrix for the beamformer. Minimum Variance Distortionless Response (MVDR) beamformer and Generalized Eigenvalue (GEV) beamformer are both investigated under the CRM-based joint training architecture. We also propose a robust mask pooling strategy among multiple channels. A long short-term memory (LSTM) based language model is utilized to re-score hypotheses which further improves the overall performance. We evaluate the proposed methods on CHiME-4 challenge dataset. The CRM based system achieves a relative 10% reduction on word error rate (WER) compared with the IRM based system. Without sequence discriminative training, our best single system already achieves an average WER 2.72% on the test set which is comparable to the state-of-the-art.

BibTeX
@inproceedings{icassp2019_jointtrainingofc,
  title = {Joint Training of Complex Ratio Mask Based Beamformer and Acoustic Model for Noise Robust Asr},
  author = {Yong Xu and Chao Weng and Like Hui and Jianming Liu and Meng Yu and Dan Su and Dong Yu},
  booktitle = {ICASSP 2019},
  year = {2019}
}
Joint Training of Complex Ratio Mask Based Beamformer and Acoustic Model for Noise Robust Asr · ICASSP 2019