ICASSP 2019accepted0 citations

Multi-speaker Emotional Acoustic Modeling for CNN-based Speech Synthesis

Heejin Choi, Sangjun Park, Jinuk Park, Minsoo Hahn

Abstract

In this paper, we investigate multi-speaker emotional acoustic modeling methods for convolutional neural network (CNN) based speech synthesis system. For emotion modeling, we extend to the speech synthesis system that learns a latent embedding space of emotion, derived from a desired emotional identity, and we use emotion code and mel-frequency spectrogram as an emotion identity. In order to model speaker variation in a text-to-speech (TTS) system, we use speaker representations such as trainable speaker embedding and speaker code. We have implemented speech synthesis systems combining speaker representation and emotion representation and compared them by experiments. Experimental results have demonstrated that the multi-speaker emotional speech synthesis approach using trainable speaker embedding and emotion representation from mel spectrogram achieves higher performance when compared with other approaches in terms of naturalness, speaker similarity, and emotion similarity.

BibTeX
@inproceedings{icassp2019_multispeakeremot,
  title = {Multi-speaker Emotional Acoustic Modeling for CNN-based Speech Synthesis},
  author = {Heejin Choi and Sangjun Park and Jinuk Park and Minsoo Hahn},
  booktitle = {ICASSP 2019},
  year = {2019}
}