← Search

Raul Fernandez

5 accepted papers

2024

Speak While You Think: Streaming Speech Synthesis During Text Generation

ICASSP 2024accepted

Large Language Models (LLMs) demonstrate impressive capabilities, yet interaction with these models is mostly facilitated through text. Using Text-To-Speech to synthesize LLM outputs typically results in notable latency, which is impractical for fluent voice conversations. We propose LLM2Speech, an…

Cited by 0SourceScholar
2021

Stable Checkpoint Selection and Evaluation in Sequence to Sequence Speech Synthesis

ICASSP 2021accepted

Autoregressive Attentive Sequence-to-Sequence (S2S) speech synthesis is considered state-of-the-art in terms of speech quality and naturalness, as evaluated on a finite set of testing utterances. However, it can occasionally suffer from stability issues at inference time, such as local intelligibili…

Cited by 0SourceScholar
2018

Measuring the Effect of Linguistic Resources on Prosody Modeling for Speech Synthesis

ICASSP 2018accepted

The generation of natural and expressive prosodic contours is an important component of a text-to-speech (TTS) system which, in most classical architectures, relies on the existence of a text-analysis processor that can extract prosody-predictive features and pass them to a statistical learning mode…

Cited by 0SourceScholar
2017

Voice-transformation-based data augmentation for prosodic classification

ICASSP 2017accepted

In this work we explore data-augmentation techniques for the task of improving the performance of a supervised recurrent-neural-network classifier tasked with predicting prosodic-boundary and pitch-accent labels. The technique is based on applying voice transformations to the training data that modi…

Cited by 12SourceScholar
2016

Using continuous lexical embeddings to improve symbolic-prosody prediction in a text-to-speech front-end

ICASSP 2016accepted

The prediction of symbolic prosodic categories from text is an important, but challenging, natural-language processing task given the various ways in which an input can be realized, and the fact that knowledge about what features determine this realization is incomplete or inaccessible to the model.…

Cited by 0SourceScholar