Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
The goal of this paper is to accelerate codec-based speech synthesis systems with minimum sacrifice to speech quality. We propose an enhanced inference method that allows for flexible trade-offs between speed and quality during inference without requiring additional training. Our core idea is to pre…