ICASSP 2019accepted0 citations

Token-wise Training for Attention Based End-to-end Speech Recognition

Peidong Wang, Jia Cui, Chao Weng, Dong Yu

Abstract

In attention based end-to-end (A-E2E) speech recognition systems, the dependency between output tokens is typically formulated as an input-output mapping in decoder. Due to such dependency, decoding errors can easily propagate along output sequence. In this paper, we propose a token-wise training (TWT) method for A-E2E models. The new method is flexible and can be combined with a variety of loss functions. Applying TWT to multiple hypotheses, we propose a novel TWT in beam (TWTiB) training scheme. Trained on the benchmark Switchboard (SWBD) 300h corpus, TWTiB outperforms the previous best training scheme on the SWBD evaluation subset.

BibTeX
@inproceedings{icassp2019_tokenwisetrainin,
  title = {Token-wise Training for Attention Based End-to-end Speech Recognition},
  author = {Peidong Wang and Jia Cui and Chao Weng and Dong Yu},
  booktitle = {ICASSP 2019},
  year = {2019}
}