Using Modified Adult Speech as Data Augmentation for Child Speech Recognition
Zijian Fan, Xinwei Cao, Giampiero Salvi, Torbjørn Svendsen
Abstract
Data augmentation is a technique which enhances the size and quality of training data such that deep learning or machine learning models can achieve better performance. This paper proposes a novel way of applying data augmentation for child speech recognition in the low data resource scenario. Data augmentation is achieved by modifying existing adult speech signals. The procedure consists of two main parts, resampling, and time scaling. The experiment involves both speech from children aged from kindergarten to grade 10, and adults’ speech. We test the proposed method using both a TDNN-HMM and a GMM-HMM acoustic model. The results show that the proposed data augmentation scheme achieves a relative 7.95% reduction of WERs compared with 4.56% relative reduction when using a traditional bilinear frequency warping approach.
BibTeX
@inproceedings{icassp2023_usingmodifiedadu,
title = {Using Modified Adult Speech as Data Augmentation for Child Speech Recognition},
author = {Zijian Fan and Xinwei Cao and Giampiero Salvi and Torbjørn Svendsen},
booktitle = {ICASSP 2023},
year = {2023}
}