Static and Dynamic State Predictions for Acoustic Model Combination
Abstract
Acoustic model combination (AMOC) is an active research area. Model combination techniques are critical for many automatic speech recognition (ASR) scenarios, and provide frameworks to combine diverse acoustic models to boost ASR performance. We scope this work in the broad framework of AMOC, and present static and dynamic state combinations of acoustic models. We motivate and rationalize the benefits from our combination techniques, and present many applications and extensions. We apply our work in the context of combining a generic and a scenario-specific (dedicated) acoustic model; we train the proposed model with an ASR objective to best align with ASR performance. We conduct our experiments on large-vocabulary ASR task with over 30k hours of training data. Compared to generic model, we demonstrate a strong 6% word error relative reduction (WERR) in average across a variety of tasks, and specifically 25% and 8% WERR for far-field speaker and an emerging car scenario.
BibTeX
@inproceedings{icassp2019_staticanddynamic,
title = {Static and Dynamic State Predictions for Acoustic Model Combination},
author = {Kshitiz Kumar and Yifan Gong},
booktitle = {ICASSP 2019},
year = {2019}
}