← Search

Manuel Sam Ribeiro

6 accepted papers

2025

Lightweight neural front-ends for low-resource on-device Text-to-Speech

ICASSP 2025accepted

We propose a lightweight neural front-end framework for on-device speech generation and highlight its benefits towards low-resource language scaling. While data-driven models have shown potential in front-end literature, especially since they can enable fast language expansion, they are often extrem…

Cited by 0SourceScholar
2022

Cross-Speaker Style Transfer for Text-to-Speech Using Data Augmentation

ICASSP 2022accepted

We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive data from a target speaker and supporting conversational expressive data from different speakers. Our goal is to build a…

Cited by 0SourceScholar
2022

Voice Filter: Few-Shot Text-to-Speech Speaker Adaptation Using Voice Conversion as a Post-Processing Module

ICASSP 2022accepted

State-of-the-art text-to-speech (TTS) systems require several hours of recorded speech data to generate high-quality synthetic speech. When using reduced amounts of training data, standard TTS models suffer from speech quality and intelligibility degradations, making training low-resource TTS system…

Cited by 0SourceScholar
2019

Speaker-independent Classification of Phonetic Segments from Raw Ultrasound in Child Speech

ICASSP 2019accepted

Ultrasound tongue imaging (UTI) provides a convenient way to visualize the vocal tract during speech production. UTI is increasingly being used for speech therapy, making it important to develop automatic methods to assist various time-consuming manual tasks currently performed by speech therapists.…

Cited by 0SourceScholar
2016

Wavelet-based decomposition of F0 as a secondary task for DNN-based speech synthesis with multi-task learning

ICASSP 2016accepted

We investigate two wavelet-based decomposition strategies of the f0 signal and their usefulness as a secondary task for speech synthesis using multi-task deep neural networks (MTL-DNN). The first decomposition strategy uses a static set of scales for all utterances in the training data. We propose a…

Cited by 0SourceScholar
2015

A multi-level representation of f0 using the continuous wavelet transform and the Discrete Cosine Transform

ICASSP 2015accepted

We propose a representation of f0 using the Continuous Wavelet Transform (CWT) and the Discrete Cosine Transform (DCT). The CWT decomposes the signal into various scales of selected frequencies, while the DCT compactly represents complex contours as a weighted sum of cosine functions. The proposed a…

Cited by 0SourceScholar