ICASSP 2016accepted0 citations

Non-negative intermediate-layer DNN adaptation for a 10-KB speaker adaptation profile

Kshitiz Kumar, Chaojun Liu, Yifan Gong

Abstract

Previously we demonstrated that speaker adaptation of acoustic models (AM) can provide significant improvement in the accuracy of large-scale speech recognition systems. In this work we discuss numerous challenges in scaling speaker adaptation to millions of speakers, where the size of speaker-dependent (SD) parameters is a critical challenge. Subsequently, we formulate an intermediate-layer adaptation framework for adaptation, upon which we build a non-negative adaptation for a very sparse set of non-negative SD parameters. We further improve this work with, (a) non-negative adaptation with a small-positive threshold, (b) setting small-positive weights in an already trained non-negative model to zero. We also discuss effective methods to store the non-negative SD parameters. We show that our methods reduce the SD parameters from 86KB for our previous best adaptation approach to 8.8KB, thus about 90% relative reduction in the size of SD parameters, and still retain 10+% word-error-rate-relative (WERR) gain over the baseline speaker-independent (SI) model.

BibTeX
@inproceedings{icassp2016_nonnegativeinter,
  title = {Non-negative intermediate-layer DNN adaptation for a 10-KB speaker adaptation profile},
  author = {Kshitiz Kumar and Chaojun Liu and Yifan Gong},
  booktitle = {ICASSP 2016},
  year = {2016}
}