Speech Separation for Low-Resource Languages
Marvin Borsdorf, Zexu Pan, Pascal Himmelmann, Haizhou Li, Tanja Schultz
Abstract
Speech separation aims to equip machines with the human ability of selective listening, i.e. to focus attention on specific information in spoken communication. Studies have shown that the language spoken in a cocktail party scenario matters. While the development of speech separation models can leverage extensive databases, for the majority of languages only very limited data is available. This work presents the very first study on speech separation for low-resource languages. We choose blind source separation as the task to be studied and analyze three strategies to overcome the data scarcity of two low-resource languages from the GlobalPhoneMS2 database. We show that data from other languages can be used to develop models that work for low-resource languages. Finetuning additionally boosts the performance, and training on multiple languages increases both performance and robustness. We show that dynamic mixing in the development helps to find a trade-off between performance and development time.
BibTeX
@inproceedings{icassp2025_speechseparation,
title = {Speech Separation for Low-Resource Languages},
author = {Marvin Borsdorf and Zexu Pan and Pascal Himmelmann and Haizhou Li and Tanja Schultz},
booktitle = {ICASSP 2025},
year = {2025}
}