ICASSP 2025accepted0 citations

Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding

Bohan Li, Hankun Wang, Situo Zhang, Yiwei Guo, Kai Yu

Abstract

The auto-regressive (AR) architecture, exemplified by models such as GPT, is extensively utilized in modern Text-to-Speech (TTS) systems. However, it often leads to considerable inference delays, primarily due to the challenges associated with next-token prediction in long speech sequences. In this work, we introduce VADUSA, one of the first approaches to accelerate AR-based TTS through speculative decoding. Our findings demonstrate that VADUSA not only delivers a significant reduction in inference time but also enhances TTS quality by employing draft heads to predict future speech tokens in an auto-regressive manner. Additionally, the incorporation of a tolerance mechanism during the sampling process further boosts performance, yielding approximately a 3× speedup in AR TTS. Moreover, our approach exhibits strong generalization across diverse datasets and various speech token types.

BibTeX
@inproceedings{icassp2025_fastandhighquali,
  title = {Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding},
  author = {Bohan Li and Hankun Wang and Situo Zhang and Yiwei Guo and Kai Yu},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding · ICASSP 2025