Quality estimation for asr k-best list rescoring in spoken language translation
Raymond W. M. Ng, Kashif Shah, Wilker Aziz, Lucia Specia, Thomas Hain
Abstract
Spoken language translation (SLT) combines automatic speech recognition (ASR) and machine translation (MT). During the decoding stage, the best hypothesis produced by the ASR system may not be the best input candidate to the MT system, but making use of multiple sub-optimal ASR results in SLT has been shown to be too complex computationally. This paper presents a method to rescore the k-best ASR output such as to improve translation quality. A translation quality estimation model is trained on a large number of features which aim to capture complementary information from both ASR and MT on translation difficulty and adequacy, as well as syntactic properties of the SLT inputs and outputs. Based on the predicted quality score, the ASR hypotheses are rescored before they are fed to the MT system. ASR confidence is found to be crucial in guiding the rescoring step. In an English-to-French speech-to-text translation task, the coupling of ASR and MT systems led to an increase of 0.5 BLEU points in translation quality.
BibTeX
@inproceedings{icassp2015_qualityestimatio,
title = {Quality estimation for asr k-best list rescoring in spoken language translation},
author = {Raymond W. M. Ng and Kashif Shah and Wilker Aziz and Lucia Specia and Thomas Hain},
booktitle = {ICASSP 2015},
year = {2015}
}