Joint acoustic factor learning for robust deep neural network based automatic speech recognition
Souvik Kundu, Gautam Mantena, Yanmin Qian, Tian Tan, Marc Delcroix, Khe Chai Sim
Abstract
Deep neural networks (DNNs) for acoustic modeling have been shown to provide impressive results on many state-of-the-art automatic speech recognition (ASR) applications. However, DNN performance degrades due to mismatches in training and testing conditions and thus adaptation is necessary. In this paper, we explore the use of discriminative auxiliary input features obtained using joint acoustic factor learning for DNN adaptation. These features are derived from a bottleneck (BN) layer of a DNN and are referred to as BN vectors. To derive these BN vectors, we explore the use of two types of joint acoustic factor learning which capture speaker and auxiliary information such as noise, phone and articulatory information of speech. In this paper, we show that these BN vectors can be used for adaptation and thereby improve the performance of an ASR system. We also show that the performance can be further improved on augmenting these BN vectors to conventional i-vectors. In this paper, experiments are performed on Aurora-4, REVERB challenge and AMI databases.
BibTeX
@inproceedings{icassp2016_jointacousticfac,
title = {Joint acoustic factor learning for robust deep neural network based automatic speech recognition},
author = {Souvik Kundu and Gautam Mantena and Yanmin Qian and Tian Tan and Marc Delcroix and Khe Chai Sim},
booktitle = {ICASSP 2016},
year = {2016}
}