← Search

Tomoki Koriyama

8 accepted papers

2023

Duration-Aware Pause Insertion Using Pre-Trained Language Model for Multi-Speaker Text-To-Speech

ICASSP 2023accepted

Pause insertion, also known as phrase break prediction and phrasing, is an essential part of TTS systems because proper pauses with natural duration significantly enhance the rhythm and intelligibility of synthetic speech. However, conventional phrasing models ignore various speakers’ different styl…

Cited by 0SourceScholar
2020

Utterance-Level Sequential Modeling for Deep Gaussian Process Based Speech Synthesis Using Simple Recurrent Unit

ICASSP 2020accepted

This paper presents a deep Gaussian process (DGP) model with a recurrent architecture for speech sequence modeling. DGP is a Bayesian deep model that can be trained effectively with the consideration of model complexity and is a kernel regression model that can have high expressibility. In the previ…

Cited by 0SourceScholar
2019

A Training Method Using DNN-guided Layerwise Pretraining for Deep Gaussian Processes

ICASSP 2019accepted

This paper proposes a novel framework of training method of deep Gaussian processes (DGPs). DGPs are deep architecture models based on stacked multiple GPs, which can overcome the limitation of single-layer GPs. Although a stochastic variational inference (SVI)- based method has been proposed for DG…

Cited by 0SourceScholar
2019

Generative Moment Matching Network-based Random Modulation Post-filter for DNN-based Singing Voice Synthesis and Neural Double-tracking

ICASSP 2019accepted

This paper proposes a generative moment matching network (GMMN)-based post-filter that provides inter-utterance pitch variation for deep neural network (DNN)-based singing voice synthesis. The natural pitch variation of a human singing voice leads to a richer musical experience and is used in double…

Cited by 0SourceScholar
2017

Duration prediction using multiple Gaussian process experts for GPR-based speech synthesis

ICASSP 2017accepted

This paper proposes an alternative multi-level approach to duration prediction for improving prosody generation in statistical parametric speech synthesis using multiple Gaussian process experts. We use two duration models at different levels, specifically, syllable and phone. First, we individually…

Cited by 7SourceScholar
2016

A speaker adaptation technique for Gaussian process regression based speech synthesis using feature space transform

ICASSP 2016accepted

In this paper, we propose a speaker adaptation technique for statistical parametric speech synthesis based on Gaussian process regression (GPR). Although it is reported that the GPR-based speech synthesis improves the naturalness of synthetic speech compared with the HMM-based speech synthesis, any…

Cited by 0SourceScholar
2015

Prosody generation using frame-based Gaussian process regression and classification for statistical parametric speech synthesis

ICASSP 2015accepted

This paper proposes novel models of F0 contours and phone durations using Gaussian process regression and classification (GPR and GPC) for statistical parametric speech synthesis. Although the use of frame-based GPR has shown the effectiveness of spectral feature modeling in previous studies, the ap…

Cited by 0SourceScholar