ICASSP 2023accepted0 citations

Speakeraugment: Data Augmentation for Generalizable Source Separation via Speaker Parameter Manipulation

Kai Wang, Yuhang Yang, Hao Huang, Ying Hu, Sheng Li

Abstract

Existing speech separation models based on deep learning typically generalize poorly due to domain mismatch. In this paper, we propose SpeakerAugment (SA), a data augmentation method for generalizable speech separation that aims to increase the diversity of speaker identity in training data, to mitigate speaker mismatch of domain mismatch. The SA consists of two sub-policies: (1) SA-Vocoder, which uses a vocoder to manipulate pitch and formants parameters of speakers. (2) SA-Spectrum, which directly performs pitch-shift and time-stretch on the spectrum of each speech signal. The SA is simple and effective. Experimental results show that using SA can significantly improve the generalization ability of models, especially for: 1) The training set with fewer speakers, e.g., WSJ0-2mix, or 2) The target test set with complex linguistic conditions, e.g., the TIMIT based test set. Moreover, as a data augmentation method, SA has good potential to be applicable to other speech related tasks. We validate this by applying SA in speech recognition, and experimental results show that the generalization ability is also improved.

BibTeX
@inproceedings{icassp2023_speakeraugmentda,
  title = {Speakeraugment: Data Augmentation for Generalizable Source Separation via Speaker Parameter Manipulation},
  author = {Kai Wang and Yuhang Yang and Hao Huang and Ying Hu and Sheng Li},
  booktitle = {ICASSP 2023},
  year = {2023}
}
Speakeraugment: Data Augmentation for Generalizable Source Separation via Speaker Parameter Manipulation · ICASSP 2023