2024
VoiceLDM: Text-to-Speech with Environmental Context
ICASSP 2024accepted
This paper presents VoiceLDM, a model designed to produce audio that accurately follows two distinct natural language text prompts: the description prompt and the content prompt. The former provides information about the overall environmental context of the audio, while the latter conveys the lingui…