← Search

Giulia Comini

4 accepted papers

2025

Lightweight neural front-ends for low-resource on-device Text-to-Speech

ICASSP 2025accepted

We propose a lightweight neural front-end framework for on-device speech generation and highlight its benefits towards low-resource language scaling. While data-driven models have shown potential in front-end literature, especially since they can enable fast language expansion, they are often extrem…

Cited by 0SourceScholar
2022

Cross-Speaker Style Transfer for Text-to-Speech Using Data Augmentation

ICASSP 2022accepted

We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive data from a target speaker and supporting conversational expressive data from different speakers. Our goal is to build a…

Cited by 0SourceScholar
2022

Voice Filter: Few-Shot Text-to-Speech Speaker Adaptation Using Voice Conversion as a Post-Processing Module

ICASSP 2022accepted

State-of-the-art text-to-speech (TTS) systems require several hours of recorded speech data to generate high-quality synthetic speech. When using reduced amounts of training data, standard TTS models suffer from speech quality and intelligibility degradations, making training low-resource TTS system…

Cited by 31SourceScholar
2021

Low-Resource Expressive Text-To-Speech Using Data Augmentation

ICASSP 2021accepted

While recent neural text-to-speech (TTS) systems perform remarkably well, they typically require a substantial amount of recordings from the target speaker reading in the desired speaking style. In this work, we present a novel 3-step methodology to circumvent the costly operation of recording large…

Cited by 0SourceScholar