Improving Multichannel Speech Recognition with Generalized Cross Correlation Inputs and Multitask Learning
Yu Zhang, Wenjie Li, Pengyuan Zhang, Yonghong Yan
Abstract
Acoustic signals from microphone arrays are used to improve performance in distant speech recognition due to the availability of spatial information. And multichannel automatic speech recognition (ASR) systems often separate speech enhancement module from acoustic modeling, which may be not optimal for improving recognition accuracy. In this work, we propose to improve multichannel speech recognition by supplying the generalized cross correlation (GCC) between microphones, which encodes spatial information, as input features to a long short-term memory (LSTM) acoustic model in parallel with the regular acoustic features. Moreover, multitask learning architecture is incorporated and shows its ability to improve the robustness of the model. We performed experiments on the AMI and ICSI meeting corpora, with results indicating that the proposed model outperforms the model trained directly on the concatenation of multiple microphone outputs and the model trained on a beamformed channel.
BibTeX
@inproceedings{icassp2018_improvingmultich,
title = {Improving Multichannel Speech Recognition with Generalized Cross Correlation Inputs and Multitask Learning},
author = {Yu Zhang and Wenjie Li and Pengyuan Zhang and Yonghong Yan},
booktitle = {ICASSP 2018},
year = {2018}
}