Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
Tokenising continuous speech into sequences of discrete tokens and modelling them with language models (LMs) has led to significant success in text-to-speech (TTS) synthesis. Despite these models can generate speech with high quality and naturalness, their synthesised samples can still suffer from a…