← Search

Jae-Sung Bae

4 accepted papers

2024

Latent Filling: Latent Space Data Augmentation for Zero-Shot Speech Synthesis

ICASSP 2024accepted

Previous works in zero-shot text-to-speech (ZS-TTS) have attempted to enhance its systems by enlarging the training data through crowd-sourcing or augmenting existing speech data. However, the use of low-quality data has led to a decline in the overall system performance. To avoid such degradation,…

Cited by 0SourceScholar
2024

Mels-Tts : Multi-Emotion Multi-Lingual Multi-Speaker Text-To-Speech System Via Disentangled Style Tokens

ICASSP 2024accepted

This paper proposes a multi-emotion, multi-lingual, and multi-speaker text-to-speech (MELS-TTS) system, employing disentangled style tokens for effective emotion transfer. In speech encompassing various attributes, such as emotional state, speaker identity, and linguistic style, disentangling these…

Cited by 0SourceScholar
2023

Avocodo: Generative Adversarial Network for Artifact-Free Vocoder

AAAI 2023technical

Neural vocoders based on the generative adversarial neural network (GAN) have been widely used due to their fast inference speed and lightweight networks while generating high-quality speech waveforms. Since the perceptually important speech components are primarily concentrated in the low-frequency…

2021

A Neural Text-to-Speech Model Utilizing Broadcast Data Mixed with Background Music

ICASSP 2021accepted

Recently, it has become easier to obtain speech data from various media such as the internet or YouTube, but directly utilizing them to train a neural text-to-speech (TTS) model is difficult. The proportion of clean speech is insufficient and the remainder includes background music. Even with the gl…

Cited by 0SourceScholar