← Search

Ji-Hyun Lee

3 accepted papers

2024

Latent Filling: Latent Space Data Augmentation for Zero-Shot Speech Synthesis

ICASSP 2024accepted

Previous works in zero-shot text-to-speech (ZS-TTS) have attempted to enhance its systems by enlarging the training data through crowd-sourcing or augmenting existing speech data. However, the use of low-quality data has led to a decline in the overall system performance. To avoid such degradation,…

Cited by 0SourceScholar
2022

HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech Synthesis

NeurIPS 2022accept

This paper presents HierSpeech, a high-quality end-to-end text-to-speech (TTS) system based on a hierarchical conditional variational autoencoder (VAE) utilizing self-supervised speech representations. Recently, single-stage TTS systems, which directly generate raw speech waveform from text, have be…

Cited by 57SourcePDFScholar
2022

PVAE-TTS: Adaptive Text-to-Speech via Progressive Style Adaptation

ICASSP 2022accepted

Adaptive text-to-speech (TTS) has attracted increasing interests for the purpose of training TTS systems without tons of high quality data. Nevertheless, existing adaptive TTS systems still show low adaptation quality for novel speakers, since it is hard to learn an extensive speaking style with lim…

Cited by 0SourceScholar