Cross-lingual lexical language discovery from audio data using multiple translations
Felix Stahlberg, Tim Schlippe, Stephan Vogel, Tanja Schultz
Abstract
Zero-resource Automatic Speech Recognition (ZR ASR) addresses target languages without given pronunciation dictionary, transcribed speech, and language model. Lexical discovery for ZR ASR aims to extract word-like chunks from speech. Lexical discovery benefits from the availability of written translations in another source language. In this paper, we improve lexical discovery even more by combining multiple source languages. We present a novel method for combining noisy word segmentations resulting in up to 11.2% relative F-score gain. When we extract word pronunciations from the combined segmentations to bootstrap an ASR system, we improve accuracy by 9.1% relative compared to the best system with only one translation, and by 50.1% compared to monolingual lexical discovery.
BibTeX
@inproceedings{icassp2015_crosslinguallexi,
title = {Cross-lingual lexical language discovery from audio data using multiple translations},
author = {Felix Stahlberg and Tim Schlippe and Stephan Vogel and Tanja Schultz},
booktitle = {ICASSP 2015},
year = {2015}
}