Speaker Adaptation For Enhancement Of Bone-Conducted Speech
Amin Edraki, Wai-Yip Chan, Jesper Jensen, Daniel Fogerty
Abstract
Deep neural network (DNN)-based speech enhancement models often face challenges in maintaining their performance for speakers not encountered during training. This challenge is exacerbated in applications such as enhancement and bandwidth extension of bone-conducted speech, where the distortion exhibits a close correlation with speaker-specific characteristics. We address this issue by introducing a bottleneck module aimed at disentangling speaker-specific characteristics from speech content in speech enhancement DNNs. A DNN model is trained for enhancement of bone-conducted speech and modified with the proposed bottleneck module. We evaluate the DNN’s adaptability to unseen speakers through fine-tuning the network with a limited amount of adaptation data. The results show that the proposed bottleneck module can enhance adaptation performance to new unseen speakers, especially when limited amount of speaker-specific adaptation data is available.
BibTeX
@inproceedings{icassp2024_speakeradaptatio,
title = {Speaker Adaptation For Enhancement Of Bone-Conducted Speech},
author = {Amin Edraki and Wai-Yip Chan and Jesper Jensen and Daniel Fogerty},
booktitle = {ICASSP 2024},
year = {2024}
}