← Search

Shinji Takaki

11 accepted papers

2023

Embedding a Differentiable Mel-Cepstral Synthesis Filter to a Neural Speech Synthesis System

ICASSP 2023accepted

This paper integrates a classic mel-cepstral synthesis filter into a modern neural speech synthesis system towards end-to-end controllable speech synthesis. Since the mel-cepstral synthesis filter is explicitly embedded in neural waveform models in the proposed system, both voice characteristics and…

Cited by 0SourceScholar
2021

Periodnet: A Non-Autoregressive Waveform Generation Model with a Structure Separating Periodic and Aperiodic Components

ICASSP 2021accepted

We propose PeriodNet, a non-autoregressive (non-AR) waveform generation model with a new model structure for modeling periodic and aperiodic components in speech waveforms. The non-AR waveform generation models can generate speech waveforms parallelly and can be used as a speech vocoder by condition…

Cited by 0SourceScholar
2020

Fast and High-Quality Singing Voice Synthesis System Based on Convolutional Neural Networks

ICASSP 2020accepted

The present paper describes singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the naturalness of synthesized singing voices. As singing voices represent a rich for…

Cited by 0SourceScholar
2020

Semi-Supervised Learning Based on Hierarchical Generative Models for End-to-End Speech Synthesis

ICASSP 2020accepted

This paper proposes a general framework of semi-supervised learning based on hierarchical generative models and adapts it to a Japanese end-to-end text-to-speech (TTS) system. In English TTS, several end-to-end systems have recently achieved sound quality close to that of natural human speech. Howev…

Cited by 0SourceScholar
2019

Investigation of Enhanced Tacotron Text-to-speech Synthesis Systems with Self-attention for Pitch Accent Language

ICASSP 2019accepted

End-to-end speech synthesis is a promising approach that directly converts raw text to speech. Although it was shown that Tacotron2 outperforms classical pipeline systems with regards to naturalness in English, its applicability to other languages is still unknown. Japanese could be one of the most…

Cited by 0SourceScholar
2019

Neural Source-filter-based Waveform Model for Statistical Parametric Speech Synthesis

ICASSP 2019accepted

Neural waveform models such as the WaveNet are used in many recent text-to-speech systems, but the original WaveNet is quite slow in waveform generation because of its autoregressive (AR) structure. Although faster non-AR models were recently reported, they may be prohibitively complicated due to th…

Cited by 0SourceScholar
2019

STFT Spectral Loss for Training a Neural Speech Waveform Model

ICASSP 2019accepted

This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural speech waveform model that predicts raw continuous speech waveform samples directly. Not only amplitude spectra but also phase spectra obtained from generated speech wav…

Cited by 26SourceScholar
2018

A Comparison of Recent Waveform Generation and Acoustic Modeling Methods for Neural-Network-Based Speech Synthesis

ICASSP 2018accepted

Recent advances in speech synthesis suggest that limitations such as the lossy nature of the amplitude spectrum with minimum phase approximation and the over-smoothing effect in acoustic modeling can be overcome by using advanced machine learning approaches. In this paper, we build a framework in wh…

Cited by 0SourceScholar
2017

Adapting and controlling DNN-based speech synthesis using input codes

ICASSP 2017accepted

Methods for adapting and controlling the characteristics of output speech are important topics in speech synthesis. In this work, we investigated the performance of DNN-based text-to-speech systems that in parallel to conventional text input also take speaker, gender, and age codes as inputs, in ord…

Cited by 0SourceScholar
2017

An autoregressive recurrent mixture density network for parametric speech synthesis

ICASSP 2017accepted

Neural-network-based generative models, such as mixture density networks, are potential solutions for speech synthesis. In this paper we follow this path and propose a recurrent mixture density network that incorporates a trainable autoregressive model. An advantage of incorporating an autoregressiv…

Cited by 0SourceScholar
2016

A deep auto-encoder based low-dimensional feature extraction from FFT spectral envelopes for statistical parametric speech synthesis

ICASSP 2016accepted

In the state-of-the-art statistical parametric speech synthesis system, a speech analysis module, e.g. STRAIGHT spectral analysis, is generally used for obtaining accurate and stable spectral envelopes, and then low-dimensional acoustic features extracted from obtained spectral envelopes are used fo…

Cited by 0SourceScholar