Probabilistic spatial dictionary based online adaptive beamforming for meeting recognition in noisy and reverberant environments
Nobutaka Ito, Shoko Araki, Marc Delcroix, Tomohiro Nakatani
Abstract
Here we propose online adaptive beamforming for automatic speech recognition (ASR) in meetings in noisy, reverberant environments. The proposed method is based on recently developed mask-based beamforming, in which accurate mask estimation and diarization are paramount. Real-world experiments have shown that mask-based beamforming enables accurate ASR in meetings in small noise and reverberation with a signal-to-noise ratio (SNR) of 15–25 dB and a reverberation time (RT) of 120–350 ms. In this paper, we deal with a more adverse condition: meetings in large noise and reverberation with an SNR of 3–15 dB and an RT of 500 ms. To this end, we exploit a probabilistic spatial dictionary, a dictionary that consists of a pre-trained probability distribution of source location features for each potential speaker location. This dictionary enables us to perform mask estimation and diarization for beamforming accurately, even in the above adverse condition. The proposed method reduced the word error rate (WER) on real meeting data by 54.8% relative to our previous beamforming method.
BibTeX
@inproceedings{icassp2017_probabilisticspa,
title = {Probabilistic spatial dictionary based online adaptive beamforming for meeting recognition in noisy and reverberant environments},
author = {Nobutaka Ito and Shoko Araki and Marc Delcroix and Tomohiro Nakatani},
booktitle = {ICASSP 2017},
year = {2017}
}