ICASSP 2017accepted0 citations

Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamforming

Tomohiro Nakatani, Nobutaka Ito, Takuya Higuchi, Shoko Araki, Keisuke Kinoshita

Abstract

Recently, time-frequency mask-based beamforming has been extensively studied as the frontend of deep neural network (DNN) based automatic speech recognition (ASR) in noisy environments. Two mask estimation approaches have been separately developed for this beamforming method, namely the the DNN-based approach, which exploits the time-frequency features of the signal, and the spatial clustering-based approach, which exploits the spatial features of the signal. This paper proposes a new method that integrates the two approaches in a probabilistic way to further improve mask estimation by exploiting the advantages of both approaches. Experiments using the real data of the CHiME-3 multichannel noisy speech corpus show that the proposed method almost always outperforms the conventional approaches in terms of word error rate (WER) improvement.

BibTeX
@inproceedings{icassp2017_integratingdnnba,
  title = {Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamforming},
  author = {Tomohiro Nakatani and Nobutaka Ito and Takuya Higuchi and Shoko Araki and Keisuke Kinoshita},
  booktitle = {ICASSP 2017},
  year = {2017}
}
Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamforming · ICASSP 2017