← Search

Yannis Stylianou

15 accepted papers

2019

An Unsupervised Learning Approach to Neural-net-supported Wpe Dereverberation

ICASSP 2019accepted

Reverberation degrades signal quality and increases word error rates in automatic speech recognition (ASR). Reverberation suppression is, thus, a key component in listening enhancement devices and ASR front end. The weighted prediction error (WPE) is a prominent and effective method that gained popu…

Cited by 0SourceScholar
2018

Adaptation of an Expressive Single Speaker Deep Neural Network Speech Synthesis System

ICASSP 2018accepted

One of the advantages of statistical parametric speech synthesis is the ability to alter some of the characteristics of the speech e.g. change the speaker, expression etc. In this paper we present a technique to adapt an expressive single speaker deep neural network (DNN) speech synthesis model to a…

Cited by 0SourceScholar
2017

Adaptive gain control and time warp for enhanced speech intelligibility under reverberation

ICASSP 2017accepted

Moderate and severe reverberation reduce speech intelligibility as a result of the overlap-masking effect, which constitutes the simultaneous observation of multiple delayed and attenuated copies of the speech signal. Recent progress has been made in ameliorating the degradation in intelligibility b…

Cited by 0SourceScholar
2017

Expressive visual text to speech and expression adaptation using deep neural networks

ICASSP 2017accepted

In this paper, we present an expressive visual text to speech system (VTTS) based on a deep neural network (DNN). Given an input text sentence and a set of expression tags, the VTTS is able to produce not only the audio speech, but also the accompanying facial movements. The expressions can either b…

Cited by 0SourceScholar
2017

Predicting dialogue success, naturalness, and length with acoustic features

ICASSP 2017accepted

Statistical methods for Spoken Dialogue Systems have been shown to reduce the cost of development, while successfully handling a variety of applications. However, such systems are usually trained with simulated users or paid subjects in controlled settings. While this may be sufficient to jump-start…

Cited by 0SourceScholar
2016

Initial investigation of speech synthesis based on complex-valued neural networks

ICASSP 2016accepted

Although frequency analysis often leads us to a speech signal in the complex domain, the acoustic models we frequently use are designed for real-valued data. Phase is usually ignored or modelled separately from spectral amplitude. Here, we propose a complex-valued neural network (CVNN) for directly…

Cited by 0SourceScholar
2016

Multi-stream spectral representation for statistical parametric speech synthesis

ICASSP 2016accepted

In statistical parametric speech synthesis such as Hidden Markov Model (HMM) based synthesis, one of the problems is in the over-smoothing of parameters, which leads to a muffled sensation in the synthesised output. In this paper, we propose an approach in which the high frequency spectrum is modell…

Cited by 0SourceScholar
2015

Improved face-to-face communication using noise reduction and speech intelligibility enhancement

ICASSP 2015accepted

Significant improvements in intelligibility of speech in noise can be obtained by modifying the speech signal in the time and/or frequency domains. However, most speech intelligibility enhancement algorithms are designed to use clean speech as an input, and their performance suffers once the input s…

Cited by 0SourceScholar
2015

Methods for applying dynamic sinusoidal models to statistical parametric speech synthesis

ICASSP 2015accepted

Sinusoidal vocoders can generate high quality speech, but they have not been extensively applied to statistical parametric speech synthesis. This paper presents two ways for using dynamic sinusoidal models for statistical speech synthesis, enabling the sinusoid parameters to be modelled in HMM-based…

Cited by 0SourceScholar
2015

Robust excitation-based features for Automatic Speech Recognition

ICASSP 2015accepted

In this paper we investigate the use of noise-robust features characterizing the speech excitation signal as complementary features to the usually considered vocal tract based features for Automatic Speech Recognition (ASR). The proposed Excitation-based Features (EBF) are tested in a state-of-the-a…

Cited by 0SourceScholar