2024
Controllable Speaking Styles Using A Large Language Model
ICASSP 2024accepted
Reference-based Text-to-Speech (TTS) models can generate multiple, prosodically-different renditions of the same target text. Such models jointly learn a latent acoustic space during training, which can be sampled from during inference. Controlling these models during inference typically requires fi…