ICASSP 2016accepted0 citations

Source modeling for HMM based speech synthesis using integrated LP residual

Nagaraj Adiga, S. R. Mahadeva Prasanna

Abstract

In this work, new method of source modeling for HMM based speech synthesis is proposed using integrated LP residual (ILPR). The nature of ILPR waveform resembles the glottal flow derivative signal and may keep the speaker characteristics in a better way. The ILPR signal is modeled in the frequency domain by dividing the spectrum into two bands to characterize harmonic and noise components of the voice speech segment. The harmonic components of ILPR signals below the maximum voiced frequency (f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">m</sub> ) is modeled using mel-cepstral coefficients called as RMCEPs, whereas noise component above f <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">m</sub> is modeled by pitch adaptive triangular noise envelope weighted by the strength of excitation (SoE). The RMCEPs and SoE are modeled on the HMM framework along with MCEPs and F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> representing vocal tract information and fundamental frequency, respectively. The synthesized speech by the proposed source modeling reduces the buzziness and improves the speaker similarity compared to the conventional impulse / noise and mixed excitation source modeling and comparable with STRAIGHT based excitation. This is further reflected in both objective and subjective valuations.

BibTeX
@inproceedings{icassp2016_sourcemodelingfo,
  title = {Source modeling for HMM based speech synthesis using integrated LP residual},
  author = {Nagaraj Adiga and S. R. Mahadeva Prasanna},
  booktitle = {ICASSP 2016},
  year = {2016}
}