WFST-based structural classification integrating dnn acoustic features and RNN language features for speech recognition
Quoc Truong Do, Satoshi Nakamura, Marc Delcroix, Takaaki Hori
Abstract
This paper proposes a method to train Weighted Finite State Transducer (WFST) based structural classifiers using deep neural network (DNN) acoustic features and recurrent neural network (RNN) language features for speech recognition. Structural classification is an effective approach to achieve highly accurate recognition of structured data in which the classifier is optimized to maximize the discriminative performance using different kinds of features. A WFST-based classifier, which can integrate acoustic, pronunciation, and language features embedded in a composed WFST, was recently extended to incorporate DNN bottleneck (DNNBN) features. In this paper, we further investigate the integration of a RNN language model (RNNLM) with the WFST classifier. To this end, we introduce a lattice rescoring method using a RNNLM for efficient classifier training. In a lecture transcription task, we reduced the word error rate from 19.2% to 18.6% by optimizing the WFST parameters for the DNNBN acoustic and RNNLM language features.
BibTeX
@inproceedings{icassp2015_wfstbasedstructu,
title = {WFST-based structural classification integrating dnn acoustic features and RNN language features for speech recognition},
author = {Quoc Truong Do and Satoshi Nakamura and Marc Delcroix and Takaaki Hori},
booktitle = {ICASSP 2015},
year = {2015}
}