Statistical Phrase/Accent Command Estimation Algorithm Utilizing Linguistic Information
Abstract
The importance of extracting non-linguistic information has been highlighted in a growing variety of applications of speech signal processing. Among the audio features carrying such information, fundamental frequency (F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> ) contours are considered primarily important. The Fujisaki model is a physical model that describes a F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> contour with only a small number of parameters, namely, the timings and magnitudes of the phrase and accent commands, and a stochastic formulation and estimation algorithm have recently been proposed for it. However, the use of linguistic information has so far been limited, while it is known that accent commands are strongly related to linguistic information in many languages, and linguistic information could be obtained from the input audio signals by using speech recognition techniques. Against this background, this paper introduces a novel F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> command parameter estimation method that incorporates linguistic information with the stochastic framework. Experiments using real speech data show that when linguistic information is appropriately utilized, the estimation accuracy of accent command parameters is improved by 43% under the proposed criteria.
BibTeX
@inproceedings{icassp2018_statisticalphras,
title = {Statistical Phrase/Accent Command Estimation Algorithm Utilizing Linguistic Information},
author = {Ryotaro Sato and Kunio Kashino},
booktitle = {ICASSP 2018},
year = {2018}
}