← Search

Lauri Juvela

15 accepted papers

2023

End-to-End Amp Modeling: from Data to Controllable Guitar Amplifier Models

ICASSP 2023accepted

This paper describes a data-driven approach to creating real-time neural network models of guitar amplifiers, recreating the amplifiers’ sonic response to arbitrary inputs at the full range of controls present on the physical device. While the focus on the paper is on the data collection pipeline, w…

Cited by 0SourceScholar
2020

Transferring Neural Speech Waveform Synthesizers to Musical Instrument Sounds Generation

ICASSP 2020accepted

Recent neural waveform synthesizers such as WaveNet, WaveG-low, and the neural-source-filter (NSF) model have shown good performance in speech synthesis despite their different methods of waveform generation. The similarity between speech and music audio synthesis techniques suggests interesting ave…

Cited by 0SourceScholar
2019

Cycle-consistent Adversarial Networks for Non-parallel Vocal Effort Based Speaking Style Conversion

ICASSP 2019accepted

Speaking style conversion (SSC) is the technology of converting natural speech signals from one style to another. In this study, we propose the use of cycle-consistent adversarial networks (CycleGANs) for converting styles with varying vocal effort, and focus on conversion between normal and Lombard…

Cited by 0SourceScholar
2019

Data Augmentation Strategies for Neural Network F0 Estimation

ICASSP 2019accepted

This study explores various speech data augmentation methods for the task of noise-robust fundamental frequency (F0) estimation with neural networks. The explored augmentation strategies are split into additive noise and channel-based augmentation and into vocoder-based augmentation methods. In voco…

Cited by 0SourceScholar
2019

Waveform Generation for Text-to-speech Synthesis Using Pitch-synchronous Multi-scale Generative Adversarial Networks

ICASSP 2019accepted

The state-of-the-art in text-to-speech (TTS) synthesis has recently improved considerably due to novel neural waveform generation methods, such as WaveNet. However, these methods suffer from their slow sequential inference process, while their parallel versions are difficult to train and even more c…

Cited by 24SourceScholar
2018

A Comparison of Recent Waveform Generation and Acoustic Modeling Methods for Neural-Network-Based Speech Synthesis

ICASSP 2018accepted

Recent advances in speech synthesis suggest that limitations such as the lossy nature of the amplitude spectrum with minimum phase approximation and the over-smoothing effect in acoustic modeling can be overcome by using advanced machine learning approaches. In this paper, we build a framework in wh…

Cited by 0SourceScholar
2018

Speech Waveform Synthesis from MFCC Sequences with Generative Adversarial Networks

ICASSP 2018accepted

This paper proposes a method for generating speech from filterbank mel frequency cepstral coefficients (MFCC), which are widely used in speech applications, such as ASR, but are generally considered unusable for speech synthesis. First, we predict fundamental frequency and voicing information from M…

Cited by 0SourceScholar
2017

Non-parallel voice conversion using i-vector PLDA: towards unifying speaker verification and transformation

ICASSP 2017accepted

Text-independent speaker verification (recognizing speakers regardless of content) and non-parallel voice conversion (transforming voice identities without requiring content-matched training utterances) are related problems. We adopt i-vector method to voice conversion. An i-vector is a fixed-dimens…

Cited by 0SourceScholar
2017

Normal-to-shouted speech spectral mapping for speaker recognition under vocal effort mismatch

ICASSP 2017accepted

Speaker recognition performance degrades substantially in case of vocal effort mismatch (e.g. shouted vs. normal speech) between test and enrollment utterances. Such a mismatch is often encountered, for example, in forensic speaker recognition. This paper introduces a novel spectral mapping method w…

Cited by 0SourceScholar
2016

High-pitched excitation generation for glottal vocoding in statistical parametric speech synthesis using a deep neural network

ICASSP 2016accepted

Achieving high quality and naturalness in statistical parametric synthesis of female voices remains to be difficult despite recent advances in the study area. Vocoding is one such key element in all statistical speech synthesizers that is known to affect the synthesis quality and naturalness. The pr…

Cited by 0SourceScholar