← Search

Yogesh Virkar

4 accepted papers

2022

Duration Modeling of Neural TTS for Automatic Dubbing

ICASSP 2022accepted

Automatic dubbing (AD) addresses the problem of translating speech in a video with speech in another language while preserving the viewer experience. A most important requirement of AD is isochrony, i.e. dubbed speech has to closely match the timing of speech and pauses of the original audio. In our…

Cited by 0SourceScholar
2022

ISOMETRIC MT: Neural Machine Translation for Automatic Dubbing

ICASSP 2022accepted

Automatic dubbing (AD) is among the machine translation (MT) use cases where translations should match a given length to allow for synchronicity between source and target speech. For neural MT, generating translations of length close to the source length (e.g. within ±10% in character count), while…

Cited by 0SourceScholar
2021

Improvements to Prosodic Alignment for Automatic Dubbing

ICASSP 2021accepted

Automatic dubbing is an extension of speech-to-speech translation such that the resulting target speech is carefully aligned in terms of duration, lip movements, timbre, emotion, prosody, etc. of the speaker in order to achieve audiovisual coherence. Dubbing quality strongly depends on isochrony, i.…

Cited by 0SourceScholar
2021

Machine Translation Verbosity Control for Automatic Dubbing

ICASSP 2021accepted

Automatic dubbing aims at seamlessly replacing the speech in a video document with synthetic speech in a different language. The task implies many challenges, one of which is generating translations that not only convey the original content, but also match the duration of the corresponding utterance…

Cited by 0SourceScholar