ICASSP 2018accepted0 citations

An Investigation of a Knowledge Distillation Method for CTC Acoustic Models

Ryoichi Takashima, Sheng Li, Hisashi Kawai

Abstract

End-to-end acoustic models, such as connectionist temporal classification (CTC) and the attention model, have been studied, and their speech recognition accuracies come close to those of conventional deep neural network (DNN)-hidden Markov models. However, most high-performance end-to-end models are not suitable for real-time (streaming) speech recognition because they are based on bidirectional recurrent neural networks (RNNs). In this study, to improve the performance of unidirectional RNN-based CTC, which is suitable for real-time processing, we investigate the knowledge distillation (KD)-based model compression method for training a CTC acoustic model. we evaluate a frame-level KD method and a sequence-level KD method for CTC model. The speech recognition experiments on Wall Street Journal tasks demonstrate that, the frame-level KD worsens the WERs ofunidirectional CTC model, whereas sequence-level KD can improve the WERs of the model.

BibTeX
@inproceedings{icassp2018_aninvestigationo,
  title = {An Investigation of a Knowledge Distillation Method for CTC Acoustic Models},
  author = {Ryoichi Takashima and Sheng Li and Hisashi Kawai},
  booktitle = {ICASSP 2018},
  year = {2018}
}
An Investigation of a Knowledge Distillation Method for CTC Acoustic Models · ICASSP 2018