2024
High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion Models
ICASSP 2024accepted
Text-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis decouples TTS by combining two types of discrete speech representations(semantic & acoustic) and using two sequence-to-seque…