ICASSP 2025accepted0 citations

Compact Neural TTS Voices for Accessibility

Kunal Jain, Eoin Murphy, Deepanshu Gupta, Jonathan Dyke, Saumya Shah, Vasilieios Tsiaras, Petko Nikolov Petkov, Alistair Conkie

Abstract

Contemporary text-to-speech solutions for accessibility applications can typically be classified into two categories: (i) device-based statistical parametric speech synthesis (SPSS) or unit selection (USEL) and (ii) cloud-based neural TTS. SPSS and USEL offer low latency and low disk footprint at the expense of naturalness and audio quality. Cloud-based neural TTS systems provide significantly better audio quality and naturalness but regress in terms of latency and responsiveness, rendering these impractical for real-world applications. More recently, neural TTS models were made de-ployable to run on handheld devices. Nevertheless, latency remains higher than SPSS and USEL, while disk footprint prohibits pre-installation for multiple voices at once. In this work, we describe a high-quality compact neural TTS system achieving latency on the order of 15 ms with low disk footprint. The proposed solution is capable of running on low-power devices.

BibTeX
@inproceedings{icassp2025_compactneuraltts,
  title = {Compact Neural TTS Voices for Accessibility},
  author = {Kunal Jain and Eoin Murphy and Deepanshu Gupta and Jonathan Dyke and Saumya Shah and Vasilieios Tsiaras and Petko Nikolov Petkov and Alistair Conkie},
  booktitle = {ICASSP 2025},
  year = {2025}
}