Cumulative moving averaged bottleneck speaker vectors for online speaker adaptation of CNN-based acoustic models
Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Taichi Asami, Shigeru Katagiri, Tomohiro Nakatani
Abstract
Adapting acoustic models to speakers have shown to greatly improve performance for many tasks. Among the adaptation approaches, exploiting auxiliary features characterizing speakers or environments has received great attention because they allow rapid adaptation, i.e. adaptation with limited amount of speech data such as a single utterance. However, the auxiliary features are usually computed in batch mode, which causes some inevitable latency. In this paper we explore an extension of the auxiliary feature-based adaptation to online processing. We employ auxiliary features obtained from bottleneck speaker vectors and extend their computation to online processing using cumulative moving averaging. We test our proposed approach for deep CNN-based acoustic models, using context adaptive networks to exploit the auxiliary features. Experimental results on the CHiME-3 task demonstrate that the proposed approach can realize online speaker adaptation.
BibTeX
@inproceedings{icassp2017_cumulativemoving,
title = {Cumulative moving averaged bottleneck speaker vectors for online speaker adaptation of CNN-based acoustic models},
author = {Tsubasa Ochiai and Marc Delcroix and Keisuke Kinoshita and Atsunori Ogawa and Taichi Asami and Shigeru Katagiri and Tomohiro Nakatani},
booktitle = {ICASSP 2017},
year = {2017}
}