Data selection for noise robust exemplar matching
Emre Yilmaz, Jort F. Gemmeke, Hugo Van hamme
Abstract
Exemplar-based acoustic modeling is based on labeled training segments that are compared with the unseen test utterances with respect to a dissimilarity measure. Using a larger number of accurately labeled exemplars provides better generalization thus improved recognition performance which comes with increased computation and memory requirements. We have recently developed a noise robust exemplar matching-based automatic speech recognition system which uses a large number of undercomplete dictionaries containing speech exemplars of the same length and label to recognize noisy speech. In this work, we investigate several speech exemplar selection techniques proposed for undercomplete speech dictionaries to find a trade-off between the recognition accuracy and the acoustic model size in terms of the amount of speech exemplars used for recognition. The exemplar selection criterion has be to chosen carefully as the amount of redundancy in these dictionaries is very limited compared to overcomplete dictionaries containing plenty of exemplars. The recognition accuracies obtained on the small vocabulary track of the 2nd CHiME Challenge and the AURORA-2 database using the complete and pruned dictionaries are compared to investigate the performance of each selection criterion.
BibTeX
@inproceedings{icassp2016_dataselectionfor,
title = {Data selection for noise robust exemplar matching},
author = {Emre Yilmaz and Jort F. Gemmeke and Hugo Van hamme},
booktitle = {ICASSP 2016},
year = {2016}
}