← Search

Oliver Watts

7 accepted papers

2026

Evaluating pretrained speech embedding systems for dysarthria detection across heterogenous datasets

ICASSP 2026poster

We present a comprehensive evaluation of pretrained speech embedding systems for the detection of dysarthric speech using existing accessible data. Dysarthric speech datasets are often small and can suffer from recording biases as well as data imbalance. To address these we selected a range of datas…

Cited by 1SourcePDFScholar
2023

PUFFIN: Pitch-Synchronous Neural Waveform Generation for Fullband Speech on Modest Devices

ICASSP 2023accepted

We present a neural vocoder designed with low-powered Alternative and Augmentative Communication devices in mind. By combining elements of successful modern vocoders with established ideas from an older generation of technology, our system is able to produce high quality synthetic speech at 48kHz on…

Cited by 0SourceScholar
2019

Speech Waveform Reconstruction Using Convolutional Neural Networks with Noise and Periodic Inputs

ICASSP 2019accepted

This paper presents a method for upsampling and transforming a compact representation of acoustics into a corresponding speech waveform. Similar to a conventional vocoder, the proposed system takes a pulse train derived from fundamental frequency and a noise sequence as inputs and shapes them to be…

Cited by 0SourceScholar
2016

From HMMS to DNNS: Where do the improvements come from?

ICASSP 2016accepted

Deep neural networks (DNNs) have recently been the focus of much text-to-speech research as a replacement for decision trees and hidden Markov models (HMMs) in statistical parametric synthesis systems. Performance improvements have been reported; however, the configuration of systems evaluated makes…

Cited by 0SourceScholar
2016

Robust TTS duration modelling using DNNS

ICASSP 2016accepted

Accurate modelling and prediction of speech-sound durations is an important component in generating more natural synthetic speech. Deep neural networks (DNNs) offer a powerful modelling paradigm, and large, found corpora of natural and expressive speech are easy to acquire for training them. Unfortu…

Cited by 0SourceScholar
2016

Wavelet-based decomposition of F0 as a secondary task for DNN-based speech synthesis with multi-task learning

ICASSP 2016accepted

We investigate two wavelet-based decomposition strategies of the f0 signal and their usefulness as a secondary task for speech synthesis using multi-task deep neural networks (MTL-DNN). The first decomposition strategy uses a static set of scales for all utterances in the training data. We propose a…

Cited by 12SourceScholar
2015

Deep neural networks employing Multi-Task Learning and stacked bottleneck features for speech synthesis

ICASSP 2015accepted

Deep neural networks (DNNs) use a cascade of hidden representations to enable the learning of complex mappings from input to output features. They are able to learn the complex mapping from text-based linguistic features to speech acoustic features, and so perform text-to-speech synthesis. Recent re…

Cited by 0SourceScholar