Dnn-Based Voice Activity Detection Using Auxiliary Speech Models in Noisy Environments
Abstract
Voice activity detection (VAD) is essential for automatic speech recognition (ASR) in noisy environments. Deep neural network (DNN)-based VAD is more powerful than previous types. In the fields of ASR and speech enhancement, to improve the performance of DNNs, in addition to spectral features, auxiliary features are used because these features are effective for adapting DNNs to a target environment. To improve the performance of DNN-based VAD further, this paper proposes two types of auxiliary feature based on auxiliary speech models. The first is activation of non-negative matrix factorization and the second is acoustic score of ASR acoustic models. These features give auxiliary information to DNNs in the same way as ASR and speech enhancement do. Experimental results for noisy VAD tasks demonstrated that DNN-based methods outperformed one of the most effective conventional methods and that both auxiliary features improved performance, with the second feature being better than the first one.
BibTeX
@inproceedings{icassp2018_dnnbasedvoiceact,
title = {Dnn-Based Voice Activity Detection Using Auxiliary Speech Models in Noisy Environments},
author = {Yuuki Tachioka},
booktitle = {ICASSP 2018},
year = {2018}
}