← Search

Toru Nakashika

4 accepted papers

2019

STFT Spectral Loss for Training a Neural Speech Waveform Model

ICASSP 2019accepted

This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural speech waveform model that predicts raw continuous speech waveform samples directly. Not only amplitude spectra but also phase spectra obtained from generated speech wav…

Cited by 26SourceScholar
2018

Parallel-Data-Free Dictionary Learning for Voice Conversion Using Non-Negative Tucker Decomposition

ICASSP 2018accepted

Voice conversion (VC) is a technique where only speaker-specific information in source speech is converted while preserving the associated phonological information. Nonnegative Matrix Factorization (NMF)-based VC has been researched because of the natural-sounding voice it produces compared with con…

Cited by 0SourceScholar
2016

Modeling deep bidirectional relationships for image classification and generation

ICASSP 2016accepted

This paper presents a novel probabilistic model that represents a joint probability of two visible variables with a deep architecture, called a deep relational model (DRM). The model stacks several layers from one visible layer on to another visible layer, sandwiching hidden layers between them. As…

Cited by 0SourceScholar
2016

Speaker adaptive model based on Boltzmann machine for non-parallel training in voice conversion

ICASSP 2016accepted

In this paper, we present a voice conversion (VC) method that does not use any parallel data while training the model. VC is a technique where only speaker specific information in source speech is converted while keeping the phonological information unchanged. Most of the existing VC methods rely on…

Cited by 0SourceScholar