ICASSP 2025accepted0 citations

Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice

Xiang Lyu, Yuxuan Wang, Tianyu Zhao, Hao Wang, Huadai Liu, Zhihao Du

Abstract

LLM-based text-to-speech(TTS) system has becoming the new trend and SOTA due to its high naturalness and zero-shot capability. However, it relies heavily on training data, usually requires at least thousands hours of labeled audio. In this report, we describe how to use pretrained CosyVoice model, to develop a streaming TTS system which supports Indian English and Indian languages. Though the pretrained CosyVoice model has never seen such data, it shows good performance in both specific speaker TTS and zero-shot voice clone after finetuning with merely 280 hours data. Experiment on LIMMITS25 challenge shows that our system achieves 4.46/4.19/4.55 naturalness, and 4.29/4.34/4.27 similarity in track1/track2/track3 respectively, which ranked 1st in all tracks.

BibTeX
@inproceedings{icassp2025_buildllmbasedzer,
  title = {Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice},
  author = {Xiang Lyu and Yuxuan Wang and Tianyu Zhao and Hao Wang and Huadai Liu and Zhihao Du},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice · ICASSP 2025