← Search

Manu Airaksinen

7 accepted papers

2019

Data Augmentation Strategies for Neural Network F0 Estimation

ICASSP 2019accepted

This study explores various speech data augmentation methods for the task of noise-robust fundamental frequency (F0) estimation with neural networks. The explored augmentation strategies are split into additive noise and channel-based augmentation and into vocoder-based augmentation methods. In voco…

Cited by 0SourceScholar
2018

Speech Waveform Synthesis from MFCC Sequences with Generative Adversarial Networks

ICASSP 2018accepted

This paper proposes a method for generating speech from filterbank mel frequency cepstral coefficients (MFCC), which are widely used in speech applications, such as ASR, but are generally considered unusable for speech synthesis. First, we predict fundamental frequency and voicing information from M…

Cited by 0SourceScholar
2017

Frequency-warped time-weighted linear prediction for glottal vocoding

ICASSP 2017accepted

Auto-regressive modeling is a prevalent source-filter separation method of speech. Conventional linear prediction (LP) and its derivatives such as weighted linear prediction (WeLP) produce parametric spectral models within a linear frequency scale, whereas frequency-warped linear prediction (WaLP) c…

Cited by 0SourceScholar
2017

Lombard speech synthesis using long short-term memory recurrent neural networks

ICASSP 2017accepted

In statistical parametric speech synthesis (SPSS), a few studies have investigated the Lombard effect, specifically by using hidden Markov model (HMM)-based systems. Recently, artificial neural networks have demonstrated promising results in SPSS, specifically by using long short-term memory recurre…

Cited by 0SourceScholar
2016

High-pitched excitation generation for glottal vocoding in statistical parametric speech synthesis using a deep neural network

ICASSP 2016accepted

Achieving high quality and naturalness in statistical parametric synthesis of female voices remains to be difficult despite recent advances in the study area. Vocoding is one such key element in all statistical speech synthesizers that is known to affect the synthesis quality and naturalness. The pr…

Cited by 0SourceScholar
2016

Quasi closed phase analysis of speech signals using time varying weighted linear prediction for accurate formant tracking

ICASSP 2016accepted

Recent research on temporally weighted linear prediction shows that quasi closed phase (QCP) analysis of speech signals provides better modeling of the vocal tract and the glottal source. Quasi closed phase analysis gives more weightage on the closed phase of the glottal cycle, at the same time deem…

Cited by 7SourceScholar