Accounting for Microprosody in Modeling Intonation
Abstract
Intonation models are often used for the generation of fundamental frequency (f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> ) contours in speech synthesis. Current intonation models only represent the intentional f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> components that are related to the phonological structure of the utterance. However, natural speech also contains non-intentional microvariations of f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> , which are usually not accounted for. Here, we derived models for two forms of microvariations: the drop in f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> during voiced obstruents, and the increased f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> at the onset of vowels following voiceless obstruents. These models were applied to remove the microvariations of f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> in a database of natural speech before the f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> contours were reproduced with the Target Approximation Model. The previously removed microvariations were then superimposed on the modeled f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> contours. The resulting model f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> contours were significantly more similar to the original (natural) f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> contours than model contours that did not account for the mi-crovariations. This approach might improve f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> modeling in future parametric speech synthesizers.
BibTeX
@inproceedings{icassp2020_accountingformic,
title = {Accounting for Microprosody in Modeling Intonation},
author = {Peter Birkholz and Xinyu Zhang},
booktitle = {ICASSP 2020},
year = {2020}
}