2024
An Experimental Comparison of Noise-Robust Text-To-Speech Synthesis Systems Based On Self-Supervised Representation
ICASSP 2024accepted
With the advance in deep learning, text-to-speech (TTS) using clean speech has witnessed significant performance improvements. As the data collected in real scenes often contain noise and thus needs to be denoised, TTS models trained on the enhanced speech suffer from distortions and residual noises…