ICASSP 2024accepted0 citations

Latent Filling: Latent Space Data Augmentation for Zero-Shot Speech Synthesis

Jae-Sung Bae, Joun Yeop Lee, Ji-Hyun Lee, Seongkyu Mun, Taehwa Kang, Hoon-Young Cho, Chanwoo Kim

Abstract

Previous works in zero-shot text-to-speech (ZS-TTS) have attempted to enhance its systems by enlarging the training data through crowd-sourcing or augmenting existing speech data. However, the use of low-quality data has led to a decline in the overall system performance. To avoid such degradation, instead of directly augmenting the input data, we propose a latent filling (LF) method that adopts simple but effective latent space data augmentation in the speaker embedding space of the ZS-TTS system. By incorporating a consistency loss, LF can be seamlessly integrated into existing ZS-TTS systems without the need for additional training stages. Experimental results show that LF significantly improves speaker similarity while preserving speech quality.

BibTeX
@inproceedings{icassp2024_latentfillinglat,
  title = {Latent Filling: Latent Space Data Augmentation for Zero-Shot Speech Synthesis},
  author = {Jae-Sung Bae and Joun Yeop Lee and Ji-Hyun Lee and Seongkyu Mun and Taehwa Kang and Hoon-Young Cho and Chanwoo Kim},
  booktitle = {ICASSP 2024},
  year = {2024}
}