2025
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
ICLR 2025poster
The recent large-scale text-to-speech (TTS) systems are usually grouped as autoregressive and non-autoregressive systems. The autoregressive systems implicitly model duration but exhibit certain deficiencies in robustness and lack of duration controllability. Non-autoregressive systems require expli…