← Search

Yoshihiko Nankaku

14 accepted papers

2024

PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model

ICASSP 2024accepted

This paper presents a neural vocoder based on a denoising diffusion probabilistic model (DDPM) incorporating explicit periodic signals as auxiliary conditioning signals. Recently, DDPM-based neural vocoders have gained prominence as non-autoregressive models that can generate high-quality waveforms.…

Cited by 0SourceScholar
2023

Embedding a Differentiable Mel-Cepstral Synthesis Filter to a Neural Speech Synthesis System

ICASSP 2023accepted

This paper integrates a classic mel-cepstral synthesis filter into a modern neural speech synthesis system towards end-to-end controllable speech synthesis. Since the mel-cepstral synthesis filter is explicitly embedded in neural waveform models in the proposed system, both voice characteristics and…

Cited by 0SourceScholar
2023

Singing Voice Synthesis Based on a Musical Note Position-Aware Attention Mechanism

ICASSP 2023accepted

This paper proposes a novel sequence-to-sequence (seq2seq) model with a musical note position-aware attention mechanism for singing voice synthesis (SVS). A seq2seq modeling approach that can simultaneously perform acoustic and temporal modeling is attractive. However, due to the difficulty of the t…

Cited by 0SourceScholar
2022

Autoregressive Variational Autoencoder with a Hidden Semi-Markov Model-Based Structured Attention for Speech Synthesis

ICASSP 2022accepted

This paper proposes an autoregressive speech synthesis model based on the variational autoencoder incorporating latent sequence representation for acoustic and linguistic features and the structure of a hidden semi-Markov model (HSMM). Although autoregressive models can provide efficient and accurat…

Cited by 0SourceScholar
2021

Periodnet: A Non-Autoregressive Waveform Generation Model with a Structure Separating Periodic and Aperiodic Components

ICASSP 2021accepted

We propose PeriodNet, a non-autoregressive (non-AR) waveform generation model with a new model structure for modeling periodic and aperiodic components in speech waveforms. The non-AR waveform generation models can generate speech waveforms parallelly and can be used as a speech vocoder by condition…

Cited by 0SourceScholar
2020

Fast and High-Quality Singing Voice Synthesis System Based on Convolutional Neural Networks

ICASSP 2020accepted

The present paper describes singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the naturalness of synthesized singing voices. As singing voices represent a rich for…

Cited by 0SourceScholar
2020

Semi-Supervised Learning Based on Hierarchical Generative Models for End-to-End Speech Synthesis

ICASSP 2020accepted

This paper proposes a general framework of semi-supervised learning based on hierarchical generative models and adapts it to a Japanese end-to-end text-to-speech (TTS) system. In English TTS, several end-to-end systems have recently achieved sound quality close to that of natural human speech. Howev…

Cited by 0SourceScholar
2019

Singing Voice Synthesis Based on Generative Adversarial Networks

ICASSP 2019accepted

This paper proposes a generative adversarial training method for deep neural network (DNN)-based singing voice synthesis. The DNN-based approach has been used in statistical parametric singing voice synthesis and improved the naturalness of the synthesized singing voice [1]. Recently, generative adv…

Cited by 0SourceScholar
2019

Speaker-dependent Wavenet-based Delay-free Adpcm Speech Coding

ICASSP 2019accepted

This paper proposes a WaveNet-based delay-free adaptive differential pulse code modulation (ADPCM) speech coding system. The WaveNet generative model, which is a state-of-the-art model for neural-network-based speech waveform synthesis, is used as the adaptive predictor in ADPCM. To further improve…

Cited by 0SourceScholar
2018

Image Recognition Based on Separable Lattice Hmms Using a Deep Neural Network for Output Probability Distributions

ICASSP 2018accepted

This paper proposes an image recognition method based on separable lattice hidden Markov models (SLHMMs) using a deep neural network (DNN) for output probability distributions. The geometric variations of the object to be recognized, e.g., size and location, are essential in image recognition. SLHMM…

Cited by 0SourceScholar
2018

Statistical Voice Conversion Based on Wavenet

ICASSP 2018accepted

This paper proposes a voice conversion technique based on WaveNet to directly generate target audio waveforms from acoustic features of a source speaker. In voice conversion based on statistical models, the relation between acoustic features, such as spectral parameters, extracted from source and ta…

Cited by 0SourceScholar
2017

Image recognition based on discriminative models using features generated from separable lattice HMMS

ICASSP 2017accepted

This paper presents an image recognition technique based on discriminative models using features generated from separable lattice hidden Markov models (SL-HMMs). A major problem in image recognition is that the recognition performance is degraded by geometric variations such as that in position and…

Cited by 0SourceScholar
2016

Trajectory training considering global variance for speech synthesis based on neural networks

ICASSP 2016accepted

This paper proposes a new training method of deep neural networks (DNNs) for statistical parametric speech synthesis. DNNs are recently used as acoustic models that represent mapping functions from linguistic features to acoustic features in statistical parametric speech synthesis. There are problem…

Cited by 0SourceScholar
2015

The effect of neural networks in statistical parametric speech synthesis

ICASSP 2015accepted

This paper investigates how to use neural networks in statistical parametric speech synthesis. Recently, deep neural networks (DNNs) have been used for statistical parametric speech synthesis. However, the specific way how DNNs should be used in statistical parametric speech synthesis has not been s…

Cited by 0SourceScholar