Enhancing noise and pitch robustness of children's ASR
Syed Shahnawazuddin, Deepak K. T., Gayadhar Pradhan, Rohit Sinha
Abstract
It is well known that, when noisy speech is transcribed using automatic speech recognition (ASR) systems trained on clean data, a highly degraded recognition performance is obtained. The problemgets further aggravatedwhen the targeted group happens to be child speakers. For children's speech, the acoustic correlates such as pitch and formant frequency vary significantly with age. This makes the recognition of children's speech very challenging. In this paper, we have explored the ways to enhance the noise robustness of ASR systems for children's speech. Towards addressing the same, recently developed front-end acoustic features based on spectral moments (SMAC) are explored. The SMAC features are reported to be more noise robust than the conventional features like the mel-frequency cepsatral coefficients. At the same time, the SMAC features are also noted to be sensitive to the variations in the pitch. To reduce the pitch sensitivity, a spectral smoothing approach based on adaptive-liftering is proposed. Spectral smoothening prior to the computation of spectral moments results in a significant improvement in the robustness to pitch without affecting the noise immunity. To further enhance noise robustness, a foreground speech segmentation and enhancement module is also included in the proposed front-end speech parameterization technique.
BibTeX
@inproceedings{icassp2017_enhancingnoisean,
title = {Enhancing noise and pitch robustness of children's ASR},
author = {Syed Shahnawazuddin and Deepak K. T. and Gayadhar Pradhan and Rohit Sinha},
booktitle = {ICASSP 2017},
year = {2017}
}