← Search

Jungil Kong

4 accepted papers

2024

ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis

ICML 2024poster

In this work, we propose a novel method for modeling numerous speakers, which enables expressing the overall characteristics of speakers in detail like a trained multi-speaker model without additional training on the target speaker's dataset. Although various works with similar purposes have been ac…

Cited by 4SourcePDFScholar
2021

Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

ICML 2021spotlight

Several recent end-to-end text-to-speech (TTS) models enabling single-stage training and parallel sampling have been proposed, but their sample quality does not match that of two-stage TTS systems. In this work, we present a parallel end-to-end TTS method that generates more natural sounding audio t…

2020

Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search

NeurIPS 2020oral

Recently, text-to-speech (TTS) models such as FastSpeech and ParaNet have been proposed to generate mel-spectrograms from text in parallel. Despite the advantage, the parallel TTS models cannot be trained without guidance from autoregressive TTS models as their external aligners. In this work, we pr…

2020

HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

NeurIPS 2020poster

Several recent work on speech synthesis have employed generative adversarial networks (GANs) to produce raw waveforms. Although such methods improve the sampling efficiency and memory usage, their sample quality has not yet reached that of autoregressive and flow-based generative models. In this wor…