← Search

Simon King

18 accepted papers

2025

Can We "Cherry-Pick"? Investigating Multiple Renditions from a Generative Speech Synthesis Model

ICASSP 2025accepted

Generative Speech Models (GSMs) have seen a surge in popularity due to their ability to generate diverse and high-quality speech. Evaluating models that generate many different renditions for a given input sentence presents a new challenge. Listening tests are still the gold standard for evaluating…

Cited by 0SourceScholar
2025

Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis

ICASSP 2025accepted

Tokenising continuous speech into sequences of discrete tokens and modelling them with language models (LMs) has led to significant success in text-to-speech (TTS) synthesis. Despite these models can generate speech with high quality and naturalness, their synthesised samples can still suffer from a…

Cited by 0SourceScholar
2023

Autovocoder: Fast Waveform Generation from a Learned Speech Representation Using Differentiable Digital Signal Processing

ICASSP 2023accepted

Most state-of-the-art Text-to-Speech systems use the mel-spectrogram as an intermediate representation, to decompose the task into acoustic modelling and waveform generation.A mel-spectrogram is extracted from the waveform by a simple, fast DSP operation, but generating a high-quality waveform from…

Cited by 0SourceScholar
2023

Ensemble Prosody Prediction For Expressive Speech Synthesis

ICASSP 2023accepted

Generating expressive speech with rich and varied prosody continues to be a challenge for Text-to-Speech. Most efforts have focused on sophisticated neural architectures intended to better model the data distribution. Yet, in evaluations it is generally found that no single model is preferred for al…

Cited by 0SourceScholar
2020

Speaker Adaptation of a Multilingual Acoustic Model for Cross-Language Synthesis

ICASSP 2020accepted

Several studies have shown promising results in adapting DNN-based acoustic models as a mechanism to transfer characteristics from pre-trained models. One such example is speaker adaptation using a small amount of data, where fine-tuning has helped train models that extrapolate well to diverse lingu…

Cited by 0SourceScholar
2019

Attentive Filtering Networks for Audio Replay Attack Detection

ICASSP 2019accepted

An attacker may use a variety of techniques to fool an automatic speaker verification system into accepting them as a genuine user. Anti-spoofing methods meanwhile aim to make the system robust against such attacks. The ASVspoof 2017 Challenge focused specifically on replay attacks, with the intenti…

Cited by 0SourceScholar
2019

Speech Waveform Reconstruction Using Convolutional Neural Networks with Noise and Periodic Inputs

ICASSP 2019accepted

This paper presents a method for upsampling and transforming a compact representation of acoustics into a corresponding speech waveform. Similar to a conventional vocoder, the proposed system takes a pulse train derived from fundamental frequency and a noise sequence as inputs and shapes them to be…

Cited by 0SourceScholar
2016

Deep neural network-guided unit selection synthesis

ICASSP 2016accepted

Vocoding of speech is a standard part of statistical parametric speech synthesis systems. It imposes an upper bound of the naturalness that can possibly be achieved. Hybrid systems using parametric models to guide the selection of natural speech units can combine the benefits of robust statistical m…

Cited by 0SourceScholar
2016

From HMMS to DNNS: Where do the improvements come from?

ICASSP 2016accepted

Deep neural networks (DNNs) have recently been the focus of much text-to-speech research as a replacement for decision trees and hidden Markov models (HMMs) in statistical parametric synthesis systems. Performance improvements have been reported; however, the configuration of systems evaluated makes…

Cited by 0SourceScholar
2016

Robust TTS duration modelling using DNNS

ICASSP 2016accepted

Accurate modelling and prediction of speech-sound durations is an important component in generating more natural synthetic speech. Deep neural networks (DNNs) offer a powerful modelling paradigm, and large, found corpora of natural and expressive speech are easy to acquire for training them. Unfortu…

Cited by 0SourceScholar
2016

Testing the consistency assumption: Pronunciation variant forced alignment in read and spontaneous speech synthesis

ICASSP 2016accepted

Forced alignment for speech synthesis traditionally aligns a phoneme sequence predetermined by the front-end text processing system. This sequence is not altered during alignment, i.e., it is forced, despite possibly being faulty. The consistency assumption is the assumption that these mistakes do n…

Cited by 14SourceScholar
2015

Attributing modelling errors in HMM synthesis by stepping gradually from natural to modelled speech

ICASSP 2015accepted

Even the best statistical parametric speech synthesis systems do not achieve the naturalness of good unit selection. We investigated possible causes of this. By constructing speech signals that lie in between natural speech and the output from a complete HMM synthesis system, we investigated various…

Cited by 0SourceScholar
2015

Deep neural networks employing Multi-Task Learning and stacked bottleneck features for speech synthesis

ICASSP 2015accepted

Deep neural networks (DNNs) use a cascade of hidden representations to enable the learning of complex mappings from input to output features. They are able to learn the complex mapping from text-based linguistic features to speech acoustic features, and so perform text-to-speech synthesis. Recent re…

Cited by 0SourceScholar
2015

SAS: A speaker verification spoofing database containing diverse attacks

ICASSP 2015accepted

This paper presents the first version of a speaker verification spoofing and anti-spoofing database, named SAS corpus. The corpus includes nine spoofing techniques, two of which are speech synthesis, and seven are voice conversion. We design two protocols, one for standard speaker verification evalu…

Cited by 0SourceScholar