ICASSP 2016accepted0 citations
Framewise speech-nonspeech classification by neural networks for voice activity detection with statistical noise suppression
Abstract
A new voice activity detection (VAD) algorithm is proposed. The proposed algorithm is the combination of augmented statistical noise suppression (ASNS) and convolutional neural network (CNN). Since the combination of ASNS and simple power thresholding was known to be very powerful for VAD under noisy conditions, even more accurate VAD is expected by replacing the power thresholding with the more elaborate classifier. Among various model-based classifiers, CNN with noise adaptive training presented the highest accuracy, and the improvement was confirmed by the experiments using CENSREC-1-C public database.
BibTeX
@inproceedings{icassp2016_framewisespeechn,
title = {Framewise speech-nonspeech classification by neural networks for voice activity detection with statistical noise suppression},
author = {Yasunari Obuchi},
booktitle = {ICASSP 2016},
year = {2016}
}