← Search

Cassia Valentini-Botinhao

7 accepted papers

2023

Autovocoder: Fast Waveform Generation from a Learned Speech Representation Using Differentiable Digital Signal Processing

ICASSP 2023accepted

Most state-of-the-art Text-to-Speech systems use the mel-spectrogram as an intermediate representation, to decompose the task into acoustic modelling and waveform generation.A mel-spectrogram is extracted from the waveform by a simple, fast DSP operation, but generating a high-quality waveform from…

Cited by 0SourceScholar
2023

Efficient Intelligibility Evaluation Using Keyword Spotting: A Study on Audio-Visual Speech Enhancement

ICASSP 2023accepted

We propose a new method for human speech intelligibility evaluation based on keyword spotting. In this method, participants play a stimulus and select the word they hear from a close set of alternatives. To find which sentence to use, the target word, and alternatives we mine a large set of stimuli…

Cited by 0SourceScholar
2023

PUFFIN: Pitch-Synchronous Neural Waveform Generation for Fullband Speech on Modest Devices

ICASSP 2023accepted

We present a neural vocoder designed with low-powered Alternative and Augmentative Communication devices in mind. By combining elements of successful modern vocoders with established ideas from an older generation of technology, our system is able to produce high quality synthetic speech at 48kHz on…

Cited by 0SourceScholar
2019

Speech Waveform Reconstruction Using Convolutional Neural Networks with Noise and Periodic Inputs

ICASSP 2019accepted

This paper presents a method for upsampling and transforming a compact representation of acoustics into a corresponding speech waveform. Similar to a conventional vocoder, the proposed system takes a pulse train derived from fundamental frequency and a noise sequence as inputs and shapes them to be…

Cited by 0SourceScholar
2016

Testing the consistency assumption: Pronunciation variant forced alignment in read and spontaneous speech synthesis

ICASSP 2016accepted

Forced alignment for speech synthesis traditionally aligns a phoneme sequence predetermined by the front-end text processing system. This sequence is not altered during alignment, i.e., it is forced, despite possibly being faulty. The consistency assumption is the assumption that these mistakes do n…

Cited by 14SourceScholar
2015

Deep neural networks employing Multi-Task Learning and stacked bottleneck features for speech synthesis

ICASSP 2015accepted

Deep neural networks (DNNs) use a cascade of hidden representations to enable the learning of complex mappings from input to output features. They are able to learn the complex mapping from text-based linguistic features to speech acoustic features, and so perform text-to-speech synthesis. Recent re…

Cited by 0SourceScholar
2015

Modelling acoustic feature dependencies with artificial neural networks: Trajectory-RNADE

ICASSP 2015accepted

Given a transcription, sampling from a good model of acoustic feature trajectories should result in plausible realizations of an utterance. However, samples from current probabilistic speech synthesis systems result in low quality synthetic speech. Henter et al. have demonstrated the need to capture…

Cited by 32SourceScholar