2026
KALL-E: Autoregressive Speech Synthesis with Next-Distribution Prediction
AAAI 2026technical
We introduce KALL-E, a novel autoregressive (AR) language model for text-to-speech (TTS) synthesis that operates by predicting the next distribution of continuous speech frames. Unlike existing methods, KALL-E directly models the continuous speech distribution conditioned on text, eliminating the ne