Robust pitch tracking in noisy speech using speaker-dependent deep neural networks
Abstract
A reliable estimate of pitch in noisy speech is crucial for many speech applications. In this paper, we propose to use speaker-dependent (SD) deep neural networks (DNNs) to model the harmonic patterns of each speaker. Specifically, SD-DNNs take spectral features as input and estimate probabilistic pitch states at each time frame. We investigate two methods for SD-DNN training. The first one is direct training when speaker-dependent data is sufficient. The second one is speaker adaptation of a speaker-independent (SI) DNN with limited data. The Viterbi algorithm is then used to track pitch through time. Experiments show that both training methods of SD-DNNs outperform an SI-DNN based system as well as a state-of-the-art pitch tracking algorithm in all SNR conditions.
BibTeX
@inproceedings{icassp2016_robustpitchtrack,
title = {Robust pitch tracking in noisy speech using speaker-dependent deep neural networks},
author = {Yuzhou Liu and DeLiang Wang},
booktitle = {ICASSP 2016},
year = {2016}
}