ICASSP 2020accepted0 citations

Accounting for Microprosody in Modeling Intonation

Peter Birkholz, Xinyu Zhang

Abstract

Intonation models are often used for the generation of fundamental frequency (f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> ) contours in speech synthesis. Current intonation models only represent the intentional f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> components that are related to the phonological structure of the utterance. However, natural speech also contains non-intentional microvariations of f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> , which are usually not accounted for. Here, we derived models for two forms of microvariations: the drop in f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> during voiced obstruents, and the increased f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> at the onset of vowels following voiceless obstruents. These models were applied to remove the microvariations of f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> in a database of natural speech before the f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> contours were reproduced with the Target Approximation Model. The previously removed microvariations were then superimposed on the modeled f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> contours. The resulting model f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> contours were significantly more similar to the original (natural) f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> contours than model contours that did not account for the mi-crovariations. This approach might improve f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> modeling in future parametric speech synthesizers.

BibTeX
@inproceedings{icassp2020_accountingformic,
  title = {Accounting for Microprosody in Modeling Intonation},
  author = {Peter Birkholz and Xinyu Zhang},
  booktitle = {ICASSP 2020},
  year = {2020}
}