ICASSP 2024accepted0 citations

Controllable Prosody Generation with Partial Inputs

Dan-Andrei Iliescu, Devang S. Ram Mohan, Tian Huey Teh, Zack Hodari

Abstract

We address the problem of human-in-the-loop control for generating prosody in the context of text-to-speech synthesis. Controlling prosody is challenging because existing generative models lack an efficient interface through which users can modify the output quickly and precisely. To solve this, we introduce a novel framework whereby the user provides partial inputs and the generative model generates the missing features. We propose a model that is specifically designed to encode partial prosodic features and output complete audio. We show empirically that our model displays two essential qualities of a human-in-the-loop control mechanism: efficiency and robustness. With even a very small number of input values (~4), our model enables users to improve the quality of the output significantly in terms of listener preference (4:1).

BibTeX
@inproceedings{icassp2024_controllablepros,
  title = {Controllable Prosody Generation with Partial Inputs},
  author = {Dan-Andrei Iliescu and Devang S. Ram Mohan and Tian Huey Teh and Zack Hodari},
  booktitle = {ICASSP 2024},
  year = {2024}
}