← Search

Taejun Bak

3 accepted papers

2024

MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech

EMNLP 2024finding

Text-to-speech (TTS) systems that scale up the amount of training data have achieved significant improvements in zero-shot speech synthesis. However, these systems have certain limitations: they require a large amount of training data, which increases costs, and often overlook prosody similarity. To…

2024

SYNTHE-SEES: Face Based Text-to-Speech for Virtual Speaker

ICASSP 2024accepted

Recent virtual voice generation researches have limitations in that they results in low-quality voice and generate inconsistent voice from the same speaker’s different facial images. To handle this, we propose a facial encoder module for the pre-trained multi-speaker TTS system called SYNTHE-SEES, w…

Cited by 0SourceScholar
2023

Avocodo: Generative Adversarial Network for Artifact-Free Vocoder

AAAI 2023technical

Neural vocoders based on the generative adversarial neural network (GAN) have been widely used due to their fast inference speed and lightweight networks while generating high-quality speech waveforms. Since the perceptually important speech components are primarily concentrated in the low-frequency…