ICASSP 2017accepted0 citations

Fast algorithm for statistical phrase/accent command estimation based on generative model incorporating spectral features

Ryotaro Sato, Hirokazu Kameoka, Kunio Kashino

Abstract

An important challenge in speech processing involves extracting non-linguistic information from a fundamental frequency (F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> ) contour of speech. We propose a fast algorithm for estimating the model parameters of the Fujisaki model, namely, the timings and magnitudes of the phrase and accent commands. Although a powerful parameter estimation framework based on a stochastic counterpart of the Fujisaki model has recently been proposed, it still had room for improvement in terms of both computational efficiency and parameter estimation accuracy. This paper describes our two contributions. First, we propose a hard expectation-maximization (EM) algorithm for parameter inference where the E step of the conventional EM algorithm is replaced with a point estimation procedure to accelerate the estimation process. Second, to improve the parameter estimation accuracy, we add a generative process of a spectral feature sequence to the generative model. This makes it possible to use linguistic or phonological information as an additional clue to estimate the timings of the accent commands. The experiments confirmed that the present algorithm was approximately 16 times faster and estimated parameters about 3% more accurately than the conventional algorithm.

BibTeX
@inproceedings{icassp2017_fastalgorithmfor,
  title = {Fast algorithm for statistical phrase/accent command estimation based on generative model incorporating spectral features},
  author = {Ryotaro Sato and Hirokazu Kameoka and Kunio Kashino},
  booktitle = {ICASSP 2017},
  year = {2017}
}