Fix it where it fails: Pronunciation learning by mining error corrections from speech logs
Zhenzhen Kou, Daisy Stanton, Fuchun Peng, Françoise Beaufays, Trevor Strohman
Abstract
The pronunciation dictionary, or lexicon, is an essential component in an automatic speech recognition (ASR) system in that incorrect pronunciations cause systematic misrecognitions. It typically consists of a list of word-pronunciation pairs written by linguists, and a grapheme-to-phoneme (G2P) engine to generate pronunciations for words not in the list. The hand-generated list can never keep pace with the growing vocabulary of a live speech recognition system, and the G2P is usually of limited accuracy. This is especially true for proper names whose pronunciations may be influenced by various historical or foreign-origin factors. In this paper, we propose a language-independent approach to detect misrecognitions and their corrections from voice search logs. We learn previously unknown pronunciations from this data, and demonstrate that they significantly improve the quality of a production-quality speech recognition system.
BibTeX
@inproceedings{icassp2015_fixitwhereitfail,
title = {Fix it where it fails: Pronunciation learning by mining error corrections from speech logs},
author = {Zhenzhen Kou and Daisy Stanton and Fuchun Peng and Françoise Beaufays and Trevor Strohman},
booktitle = {ICASSP 2015},
year = {2015}
}