An estimation method of voice timbre evaluation values using feature extraction with Gaussian mixture model based on reference singer
Soichi Yamane, Kazuhiro Kobayashi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Satoshi Nakamura
Abstract
This paper presents an estimation method of voice timbre evaluation values for arbitrary singer's singing voices generated with a singing voice synthesis system towards the development of a singing voice retrieval system. The voice timbre evaluation values are numerical values corresponding to voice timbre expression words, such as "Age" and "Gender", and they usually need to be manually assigned to individual singers' singing voices through listening. To make it possible to automatically estimate them from given singer's singing voices, an acoustic feature to well capture only each singer's voice timbre is extracted with a Gaussian mixture model trained using parallel data between singing voices sung by many pre-stored target singers and same voices sung by a reference singer. Then, the voice timbre evaluation values are estimated from the extracted feature using regression models. The experimental results showed that the proposed method is capable of accurately estimating those values for some expression words, such as "Age" and "Gender", and nonlinear regression is effective for the expression words, "Powerfulness" and "Uniqueness."
BibTeX
@inproceedings{icassp2016_anestimationmeth,
title = {An estimation method of voice timbre evaluation values using feature extraction with Gaussian mixture model based on reference singer},
author = {Soichi Yamane and Kazuhiro Kobayashi and Tomoki Toda and Tomoyasu Nakano and Masataka Goto and Satoshi Nakamura},
booktitle = {ICASSP 2016},
year = {2016}
}