← Search

Paavo Alku

18 accepted papers

2025

Wavelet Scattering Network Features for Intensity Category Classification and Prediction of SPL from Speech

ICASSP 2025accepted

Speakers change vocal intensity in daily life to communicate over long distances and to express vocal emotions. Humans produce speech using different intensity categories (e.g. soft, normal and loud voice) and they can regulate intensity across a wide sound pressure level (SPL) range. Knowing the in…

Cited by 0SourceScholar
2023

Automatic Classification of Vocal Intensity Category from Speech

ICASSP 2023accepted

Regulation of vocal intensity is a fundamental phenomenon in speech communication. Vocal intensity can be quantified using sound pressure level (SPL), which can be measured easily by recording a standard calibration signal with speech and by comparing the energy of the recorded speech signal with th…

Cited by 0SourceScholar
2023

Utilizing Wav2Vec In Database-Independent Voice Disorder Detection

ICASSP 2023accepted

Automatic detection of voice disorders from acoustic speech signals can help to improve reliability of medical diagnosis. However, the real-life environment in which speech signals are recorded for diagnosis can be different from the environment in which the detection system’s training data was orig…

Cited by 0SourceScholar
2023

Wav2vec-Based Detection and Severity Level Classification of Dysarthria From Speech

ICASSP 2023accepted

Automatic detection and severity level classification of dysarthria directly from acoustic speech signals can be used as a tool in medical diagnosis. In this work, the pre-trained wav2vec 2.0 model is studied as a feature extractor to build detection and severity level classification systems for dys…

Cited by 0SourceScholar
2020

Comparison of Glottal Closure Instants Detection Algorithms for Emotional Speech

ICASSP 2020accepted

In production of voiced speech, epochs or glottal closure instants (GCIs) refer to the instants of significant excitation of the vocal tract. Extraction of GCIs is used as a pre-processing stage in many areas of speech technology, such as in prosody modification, speech synthesis and voice source an…

Cited by 10SourceScholar
2020

Study of Formant Modification for Children ASR

ICASSP 2020accepted

The performance of automatic speech recognition systems for children’s speech is known to suffer from the large variation and mismatch in the acoustic and linguistic attributes between children’s and adults’ speech. One of the various identified sources of mismatch is the difference in formant frequ…

Cited by 0SourceScholar
2019

Cycle-consistent Adversarial Networks for Non-parallel Vocal Effort Based Speaking Style Conversion

ICASSP 2019accepted

Speaking style conversion (SSC) is the technology of converting natural speech signals from one style to another. In this study, we propose the use of cycle-consistent adversarial networks (CycleGANs) for converting styles with varying vocal effort, and focus on conversion between normal and Lombard…

Cited by 16SourceScholar
2019

Data Augmentation Strategies for Neural Network F0 Estimation

ICASSP 2019accepted

This study explores various speech data augmentation methods for the task of noise-robust fundamental frequency (F0) estimation with neural networks. The explored augmentation strategies are split into additive noise and channel-based augmentation and into vocoder-based augmentation methods. In voco…

Cited by 0SourceScholar
2019

Waveform Generation for Text-to-speech Synthesis Using Pitch-synchronous Multi-scale Generative Adversarial Networks

ICASSP 2019accepted

The state-of-the-art in text-to-speech (TTS) synthesis has recently improved considerably due to novel neural waveform generation methods, such as WaveNet. However, these methods suffer from their slow sequential inference process, while their parallel versions are difficult to train and even more c…

Cited by 24SourceScholar
2018

Speech Waveform Synthesis from MFCC Sequences with Generative Adversarial Networks

ICASSP 2018accepted

This paper proposes a method for generating speech from filterbank mel frequency cepstral coefficients (MFCC), which are widely used in speech applications, such as ASR, but are generally considered unusable for speech synthesis. First, we predict fundamental frequency and voicing information from M…

Cited by 0SourceScholar
2017

Frequency-warped time-weighted linear prediction for glottal vocoding

ICASSP 2017accepted

Auto-regressive modeling is a prevalent source-filter separation method of speech. Conventional linear prediction (LP) and its derivatives such as weighted linear prediction (WeLP) produce parametric spectral models within a linear frequency scale, whereas frequency-warped linear prediction (WaLP) c…

Cited by 0SourceScholar
2017

Lombard speech synthesis using long short-term memory recurrent neural networks

ICASSP 2017accepted

In statistical parametric speech synthesis (SPSS), a few studies have investigated the Lombard effect, specifically by using hidden Markov model (HMM)-based systems. Recently, artificial neural networks have demonstrated promising results in SPSS, specifically by using long short-term memory recurre…

Cited by 0SourceScholar
2017

Non-parallel voice conversion using i-vector PLDA: towards unifying speaker verification and transformation

ICASSP 2017accepted

Text-independent speaker verification (recognizing speakers regardless of content) and non-parallel voice conversion (transforming voice identities without requiring content-matched training utterances) are related problems. We adopt i-vector method to voice conversion. An i-vector is a fixed-dimens…

Cited by 0SourceScholar
2017

Normal-to-shouted speech spectral mapping for speaker recognition under vocal effort mismatch

ICASSP 2017accepted

Speaker recognition performance degrades substantially in case of vocal effort mismatch (e.g. shouted vs. normal speech) between test and enrollment utterances. Such a mismatch is often encountered, for example, in forensic speaker recognition. This paper introduces a novel spectral mapping method w…

Cited by 6SourceScholar
2016

A subjective listening test of six different artificial bandwidth extension approaches in English, Chinese, German, and Korean

ICASSP 2016accepted

In studies on artificial bandwidth extension (ABE), there is a lack of international coordination in subjective tests between multiple methods and languages. Here we present the design of absolute category rating listening tests evaluating 12 ABE variants of six approaches in multiple languages, nam…

Cited by 21SourceScholar
2016

High-pitched excitation generation for glottal vocoding in statistical parametric speech synthesis using a deep neural network

ICASSP 2016accepted

Achieving high quality and naturalness in statistical parametric synthesis of female voices remains to be difficult despite recent advances in the study area. Vocoding is one such key element in all statistical speech synthesizers that is known to affect the synthesis quality and naturalness. The pr…

Cited by 0SourceScholar
2016

Quasi closed phase analysis of speech signals using time varying weighted linear prediction for accurate formant tracking

ICASSP 2016accepted

Recent research on temporally weighted linear prediction shows that quasi closed phase (QCP) analysis of speech signals provides better modeling of the vocal tract and the glottal source. Quasi closed phase analysis gives more weightage on the closed phase of the glottal cycle, at the same time deem…

Cited by 7SourceScholar