← Search

Young-Sun Joo

5 accepted papers

2024

MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech

EMNLP 2024finding

Text-to-speech (TTS) systems that scale up the amount of training data have achieved significant improvements in zero-shot speech synthesis. However, these systems have certain limitations: they require a large amount of training data, which increases costs, and often overlook prosody similarity. To…

2024

SYNTHE-SEES: Face Based Text-to-Speech for Virtual Speaker

ICASSP 2024accepted

Recent virtual voice generation researches have limitations in that they results in low-quality voice and generate inconsistent voice from the same speaker’s different facial images. To handle this, we propose a facial encoder module for the pre-trained multi-speaker TTS system called SYNTHE-SEES, w…

Cited by 0SourceScholar
2023

Avocodo: Generative Adversarial Network for Artifact-Free Vocoder

AAAI 2023technical

Neural vocoders based on the generative adversarial neural network (GAN) have been widely used due to their fast inference speed and lightweight networks while generating high-quality speech waveforms. Since the perceptually important speech components are primarily concentrated in the low-frequency…

2021

A Neural Text-to-Speech Model Utilizing Broadcast Data Mixed with Background Music

ICASSP 2021accepted

Recently, it has become easier to obtain speech data from various media such as the internet or YouTube, but directly utilizing them to train a neural text-to-speech (TTS) model is difficult. The proportion of clean speech is insufficient and the remainder includes background music. Even with the gl…

Cited by 0SourceScholar
2015

Improved time-frequency trajectory excitation modeling for a statistical parametric speech synthesis system

ICASSP 2015accepted

This paper proposes an improved time-frequency trajectory excitation (TFTE) modeling method for a statistical parametric speech synthesis system. The proposed approach overcomes the dimensional variation problem of the training process caused by the inherent nature of the pitch-dependent analysis pa…

Cited by 0SourceScholar