Sequence training of multi-task acoustic models using meta-state labels
Abstract
In this paper, we describe a multi-task learning approach for acoustic modeling where the multiple output layers are used to predict context-dependent (CD) states from different state inventories. Unlike the traditional multitask learning approach which defines a primary and secondary output layers but discards the secondary output after training, we propose to use all output layers for recognition. This can be achieved by designing a decoding network operating on tuples of CD states and combining the scores of the different outputs during search. To support training such models using a sequence-based criterion, we propose to replace the multiple output layers with a single layer encoding the CD state tuples as "meta-states". Experimental results are given on a large Voice Search task evaluated on children's speech.
BibTeX
@inproceedings{icassp2016_sequencetraining,
title = {Sequence training of multi-task acoustic models using meta-state labels},
author = {Olivier Siohan},
booktitle = {ICASSP 2016},
year = {2016}
}