Online environmental adaptation of CNN-based acoustic models using spatial diffuseness features
Christian Huemmer, Marc Delcroix, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Walter Kellermann
Abstract
We propose a new concept for adapting CNN-based acoustic models using spatial diffuseness features as auxiliary information about the acoustic environment: the spatial diffuseness features are simultaneously employed as acoustic-model input features and to estimate environmental cues for context adaptation, where one convolutional layer is factorized into several sub-layers to represent different acoustic conditions. This context-adaptive CNN-based acoustic model facilitates an online environmental adaptation and is experimentally verified for the real-world recordings provided by the CHiME-3 task. The best performing setup reduces the average word error rate scores achieved by the baseline system (without using spatial diffuseness features) from 19.4% to 15.9% and 12.2% to 10.7% considering two experimental setups with and without front-end signal enhancement, respectively.
BibTeX
@inproceedings{icassp2017_onlineenvironmen,
title = {Online environmental adaptation of CNN-based acoustic models using spatial diffuseness features},
author = {Christian Huemmer and Marc Delcroix and Atsunori Ogawa and Keisuke Kinoshita and Tomohiro Nakatani and Walter Kellermann},
booktitle = {ICASSP 2017},
year = {2017}
}