ICASSP 2016accepted0 citations

Novel neural network based fusion for multistream ASR

Sri Harish Reddy Mallidi, Hynek Hermansky

Abstract

Robustness of automatic speech recognition (ASR) to acoustic mismatches can be improved by multistream framework. Frequently used approach to combine decisions from individual streams involve training large number of neural networks, one for each possible stream combination. In this work, we propose to simplify the fusion by replacing the large number of fusion networks with a single fusion network. During training of the proposed fusion network, features from a stream are randomly dropped out. At test time, corrupted streams are identified and dropped out to improve robustness. Using the proposed approach, we were able to achieve significant reduction in number of parameters, while remaining in less than 2.5 % relative degradation of conventional fusion technique. Furthermore, proposed fusion network is also applied in a multistream ASR system to improve noise robustness of Aurora4 speech recognition task. Noticeable improvements were observed over baseline systems (relative improvement of 9.2 % in microphone mismatch and 3.2 % in additive noise conditions).

BibTeX
@inproceedings{icassp2016_novelneuralnetwo,
  title = {Novel neural network based fusion for multistream ASR},
  author = {Sri Harish Reddy Mallidi and Hynek Hermansky},
  booktitle = {ICASSP 2016},
  year = {2016}
}
Novel neural network based fusion for multistream ASR · ICASSP 2016