2022
More Than Words: In-the-Wild Visually-Driven Prosody for Text-to-Speech
CVPR 2022poster
In this paper we present VDTTS, a Visually-Driven Text-to-Speech model. Motivated by dubbing, VDTTS takes advantage of video frames as an additional input alongside text, and generates speech that matches the video signal. We demonstrate how this allows VDTTS to, unlike plain TTS models, generate sp…