Incremental Semi-Supervised Learning for Multi-Genre Speech Recognition
Banriskhem K. Khonglah, Srikanth R. Madikeri, Subhadeep Dey, Hervé Bourlard, Petr Motlícek, Jayadev Billa
Abstract
In this work, we explore a data scheduling strategy for semi-supervised learning (SSL) for acoustic modeling in automatic speech recognition. The conventional approach uses a seed model trained with supervised data to automatically recognize the entire set of unlabeled (auxiliary) data to generate new labels for subsequent acoustic model training. In this paper, we propose an approach in which the unlabelled set is divided into multiple equal-sized subsets. These subsets are processed in an incremental fashion: for each iteration a new subset is added to the data used for SSL, starting from only one subset in the first iteration. The acoustic model from the previous iteration becomes the seed model for the next one. This scheduling strategy is compared to the approach employing all unlabeled data in one-shot for training. Experiments using lattice-free maximum mutual information based acoustic model training on Fisher English gives 80% word error recovery rate. On the multi-genre evaluation sets on Lithuanian and Bulgarian relative improvements of up to 17.2% in word error rate are observed.
BibTeX
@inproceedings{icassp2020_incrementalsemis,
title = {Incremental Semi-Supervised Learning for Multi-Genre Speech Recognition},
author = {Banriskhem K. Khonglah and Srikanth R. Madikeri and Subhadeep Dey and Hervé Bourlard and Petr Motlícek and Jayadev Billa},
booktitle = {ICASSP 2020},
year = {2020}
}