Discrete Unit-based Low-latency Multi-lingual Speech Synthesis for LIMMITS'25 Challenge
Yu Jiang, Cheng Gong, Tianrui Wang, Chunyu Qiang, Haoyu Wang, Qiuyu Liu, Yuheng Lu, Xiaobao Wang
Abstract
In this paper, we present the system developed by our team, CCATTS, for the LIMMITS’25 challenge, focusing on few-shot and zero-shot TTS. We adopt a two-stage TTS strategy. In track 1, we fine-tune the pre-trained ZMM-TTS model and successfully achieve multilingual low-latency TTS. In track 2, we propose a token-based framework by modifying the first-stage model of ZMM-TTS, disentangling speech features into four types of discrete tokens—Content, Acoustic, Emotion, and Speaker—and integrating it with our designed token2wav module. This module consists of a HiFi-GAN-style decoder, an acoustic refiner, and U-Net flow matching, to generate high-quality speech. The official competition results demonstrate that our method achieves strong performance in both tracks.
BibTeX
@inproceedings{icassp2025_discreteunitbase,
title = {Discrete Unit-based Low-latency Multi-lingual Speech Synthesis for LIMMITS'25 Challenge},
author = {Yu Jiang and Cheng Gong and Tianrui Wang and Chunyu Qiang and Haoyu Wang and Qiuyu Liu and Yuheng Lu and Xiaobao Wang and Xiaolei Zhang and Longbiao Wang and Jianwu Dang},
booktitle = {ICASSP 2025},
year = {2025}
}