ICASSP 2025accepted0 citations

Stream-TTS: A Low-Latency Text-to-Speech using Kolmogorov-Arnold Networks for Streaming Speech Applications

Giridhar Pamisetty, Riya Ann Easow, Kaustubh Gupta, K. Sri Rama Murty

Abstract

The rise of conversational AI and multimodal streaming applications has led to a significant demand for low-latency Text-to-Speech (TTS) systems. This work presents a multilingual low-latency model that leverages the functional decomposition principles of Kolmogorov-Arnold Networks (KANs) in modeling the nonlinear functions with lesser parameters and smaller computational graphs than multi-layer perceptrons (MLPs). To enhance the learning capabilities of the proposed compact model, we use a supervised auxiliary learning approach with multiple subtasks whose target labels are extracted from the speech signal. We use a multispeaker nonlinear vocoder to reconstruct the natural-sounding speech from the Melspectrograms. The proposed model achieved a very low latency with a real-time factor (RTF) of 0.0795 and a mean opinion score (MOS) of 4.09 for naturalness in the multilingual streaming TTS challenge organized as part of ICASSP-2025.

BibTeX
@inproceedings{icassp2025_streamttsalowlat,
  title = {Stream-TTS: A Low-Latency Text-to-Speech using Kolmogorov-Arnold Networks for Streaming Speech Applications},
  author = {Giridhar Pamisetty and Riya Ann Easow and Kaustubh Gupta and K. Sri Rama Murty},
  booktitle = {ICASSP 2025},
  year = {2025}
}