Analysis of natural and synthetic speech using Fujisaki model
Tanvina B. Patel, Hemant A. Patil
Abstract
Text-to-speech (TTS) synthesis systems are being advanced to achieve naturalness and intelligibility in synthetic speech. Unit selection-based synthesis (USS) and Hidden Markov Model-based text-to-speech synthesis systems (HTS) are recent techniques in this area. USS-based synthetic speech is known to be natural (due to concatenation of natural speech sound units). On the other hand, HTS-based speech is not as natural in perception as USS-based synthetic speech. Due to speech synthesis technologies, voice biometrics systems may face threats due to impostor attacks. Thus, it is important to study the differences that exist between natural and synthetic speech. In this context, we investigate the effectiveness of parameters of Fujisaki model for capturing Fundamental frequency (F0) contour variations in natural and synthetic speech. F0 contour of speech contains linguistic and non-linguistic information. Experimental results on several utterances from Gujarati (a low resourced language) demonstrate the effectiveness of phrase and accent components to analyze the difference between these two speeches. Variability in phrase and accent components suggests that synthetic speech differs in terms of prosodic information in excitation source as compared to natural speech. These findings may assist to distinguish these two speeches and provide an aid to alleviate impostor attacks.
BibTeX
@inproceedings{icassp2016_analysisofnatura,
title = {Analysis of natural and synthetic speech using Fujisaki model},
author = {Tanvina B. Patel and Hemant A. Patil},
booktitle = {ICASSP 2016},
year = {2016}
}