← Search

Miaosen Wang

1 accepted papers

2022

More Than Words: In-the-Wild Visually-Driven Prosody for Text-to-Speech

CVPR 2022poster

In this paper we present VDTTS, a Visually-Driven Text-to-Speech model. Motivated by dubbing, VDTTS takes advantage of video frames as an additional input alongside text, and generates speech that matches the video signal. We demonstrate how this allows VDTTS to, unlike plain TTS models, generate sp…

Cited by 21PDFScholar