← Search

Yannis Agiomyrgiannakis

5 accepted papers

2018

B-Spline Pdf: A Generalization of Histograms to Continuous Density Models for Generative Audio Networks

ICASSP 2018accepted

Many modern neural networks use histograms to efficiently model continuous random variables. This implies that the parametric space of the multinomial distribution is easier for training large neural networks. In applications like generative audio networks, this approach introduces audible quantizat…

Cited by 0SourceScholar
2018

Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions

ICASSP 2018accepted

This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by a modified WaveNet model acting as a voc…

Cited by 0SourceScholar
2016

The matching-minimization algorithm, the INCA algorithm and a mathematical framework for voice conversion with unaligned corpora

ICASSP 2016accepted

This paper presents a mathematical framework that is suitable for voice conversion and adaptation in speech processing. Voice conversion is formulated as a search for the optimal correspondances between a set of source-speaker spectra and a set of target-speaker spectra under a transform that compen…

Cited by 0SourceScholar
2016

Voice Morphing that improves TTS quality using an optimal dynamic frequency warping-and-weighting transform

ICASSP 2016accepted

Dynamic Frequency Warping (DFW) is widely used to align spectra of different speakers. It has long been argued that frequency warping captures inter-speaker differences but DFW practice always involves a tricky preprocessing part to remove spectral tilt. The DFW residual is successfully used in Voic…

Cited by 0SourceScholar