ICASSP 2025accepted0 citations

DecoupledSynth: Enhancing Zero-Shot Text-to-Speech Via Factors Decoupling

Jingyuan Xing, Shuaiqi Chen, Xiangmin Xu, Xiaofen Xing

Abstract

Studies of speech representation enhance zero-shot Text-to-Speech by mapping text to intermediate representations before generating speech. However, using representations often struggles to balance linguistic, para-linguistic, and non-linguistic information in speech during the synthesis phase. Additionally, it usually takes substantial resources for representation extraction training. To address these limitations, we propose DecoupledSynth. It combines different self-supervised models to extract comprehensive, decoupled representations. This structure enables more thorough and nuanced synthesis by leveraging reference speech with decoupled processing stages. Experiments on the VCTK and LibriTTS datasets support the potential of this new framework, showing that it can produce more consistent and realistic speech. Speech demos are available at https://test1634.github.io/DecoupledSynth/.

BibTeX
@inproceedings{icassp2025_decoupledsynthen,
  title = {DecoupledSynth: Enhancing Zero-Shot Text-to-Speech Via Factors Decoupling},
  author = {Jingyuan Xing and Shuaiqi Chen and Xiangmin Xu and Xiaofen Xing},
  booktitle = {ICASSP 2025},
  year = {2025}
}