ICASSP 2018accepted0 citations

Parallel-Data-Free Dictionary Learning for Voice Conversion Using Non-Negative Tucker Decomposition

Yuki Takashima, Hajime Yano, Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki

Abstract

Voice conversion (VC) is a technique where only speaker-specific information in source speech is converted while preserving the associated phonological information. Nonnegative Matrix Factorization (NMF)-based VC has been researched because of the natural-sounding voice it produces compared with conventional Gaussian Mixture Model-based VC. In conventional NMF- VC, parallel data are used to train the models; therefore, unnatural pre-processing of speech data to make parallel data is needed. NMF-VC also tends to be a large model because this method has many parallel exemplars for the dictionary matrix; therefore, the computational cost is high. In this paper, we propose a novel parallel dictionary learning method using non-negative Tucker decomposition (NTD) which uses tensor decomposition and decomposes an input observation into a set of mode matrices and one core tensor. Our proposed NTD-based dictionary learning method estimates the dictionary matrix for NMF- VC without using parallel data. Experimental results show that our proposed method outperforms conventional non-parallel VC methods.

BibTeX
@inproceedings{icassp2018_paralleldatafree,
  title = {Parallel-Data-Free Dictionary Learning for Voice Conversion Using Non-Negative Tucker Decomposition},
  author = {Yuki Takashima and Hajime Yano and Toru Nakashika and Tetsuya Takiguchi and Yasuo Ariki},
  booktitle = {ICASSP 2018},
  year = {2018}
}