← Search

Roberto Barra-Chicote

8 accepted papers

2022

Duration Modeling of Neural TTS for Automatic Dubbing

ICASSP 2022accepted

Automatic dubbing (AD) addresses the problem of translating speech in a video with speech in another language while preserving the viewer experience. A most important requirement of AD is isochrony, i.e. dubbed speech has to closely match the timing of speech and pauses of the original audio. In our…

Cited by 0SourceScholar
2022

Text-Free Non-Parallel Many-To-Many Voice Conversion Using Normalising Flow

ICASSP 2022accepted

Non-parallel voice conversion (VC) is typically achieved using lossy representations of the source speech. However, ensuring only speaker identity information is dropped whilst all other information from the source speech is retained is a large challenge. This is particularly challenging in the scen…

Cited by 0SourceScholar
2022

Voice Filter: Few-Shot Text-to-Speech Speaker Adaptation Using Voice Conversion as a Post-Processing Module

ICASSP 2022accepted

State-of-the-art text-to-speech (TTS) systems require several hours of recorded speech data to generate high-quality synthetic speech. When using reduced amounts of training data, standard TTS models suffer from speech quality and intelligibility degradations, making training low-resource TTS system…

Cited by 0SourceScholar
2021

Exploring the application of synthetic audio in training keyword spotters

ICASSP 2021accepted

The study of keyword spotting, a subfield within the broader field of speech recognition that centers around identifying individual keywords in speech audio, has gained particular importance in recent years with the rise of personal voice assistants such as Alexa. As voice assistants aim to rapidly…

Cited by 0SourceScholar
2021

Improvements to Prosodic Alignment for Automatic Dubbing

ICASSP 2021accepted

Automatic dubbing is an extension of speech-to-speech translation such that the resulting target speech is carefully aligned in terms of duration, lip movements, timbre, emotion, prosody, etc. of the speaker in order to achieve audiovisual coherence. Dubbing quality strongly depends on isochrony, i.…

Cited by 0SourceScholar
2021

Machine Translation Verbosity Control for Automatic Dubbing

ICASSP 2021accepted

Automatic dubbing aims at seamlessly replacing the speech in a video document with synthetic speech in a different language. The task implies many challenges, one of which is generating translations that not only convey the original content, but also match the duration of the corresponding utterance…

Cited by 0SourceScholar
2020

BOFFIN TTS: Few-Shot Speaker Adaptation by Bayesian Optimization

ICASSP 2020accepted

We present BOFFIN TTS (Bayesian Optimization For FIne-tuning Neural Text To Speech), a novel approach for few-shot speaker adaptation. Here, the task is to fine-tune a pre-trained TTS model to mimic a new speaker using a small corpus of target utterances. We demonstrate that there does not exist a o…

Cited by 0SourceScholar
2020

Using Vaes and Normalizing Flows for One-Shot Text-To-Speech Synthesis of Expressive Speech

ICASSP 2020accepted

We propose a Text-to-Speech method to create an unseen expressive style using one utterance of expressive speech of around one second. Specifically, we enhance the disentanglement capabilities of a state-of-the-art sequence-to-sequence based system with a Variational AutoEncoder (VAE) and a Househol…

Cited by 33SourceScholar