Augmenting Short Enrollment Speech via Synthesis for Target Speaker Extraction
Zikang Huang, Jingru Lin, Meng Ge, Yu Jiang, Xiaobao Wang, Longbiao Wang, Jianwu Dang
Abstract
A high-quality enrollment speech is crucial to target speaker extraction (TSE), since it provides essential cues for identifying the target speaker in the mixture. However, real applications usually only permit a short enrollment speech, e.g. a wakeup word for a mobile device, that provides limited cues. To address this issue, we propose an enrollment augmentation strategy that allows us to enrich the limited enrollment speech with massive text data through speech synthesis. By doing so, the extended enrollment speech contains enhanced speaker timbre and phonetic content which leads to better extraction quality. Furthermore, we propose a training data augmentation strategy to improve the model’s robustness and generalization in short enrollment speech scenarios. Experiments on Libri2Mix demonstrate that our proposed strategies bring a significant improvement in extreme scenarios where only 0.5s and 1-word enrollment speech is provided. We also release our code at https://github.com/HuangZikang-TJU/Aug4TSE.
BibTeX
@inproceedings{icassp2025_augmentingshorte,
title = {Augmenting Short Enrollment Speech via Synthesis for Target Speaker Extraction},
author = {Zikang Huang and Jingru Lin and Meng Ge and Yu Jiang and Xiaobao Wang and Longbiao Wang and Jianwu Dang},
booktitle = {ICASSP 2025},
year = {2025}
}