ICASSP 2024accepted0 citations

Dialog Modeling in Audiobook Synthesis

Cheng-chieh Yeh, Amirreza Shirani, Weicheng Zhang, Tuomo Raitio, Ramya Rasipuram, Ladan Golipour, David Winarsky

Abstract

In audiobook synthesis, it is important to have the ability to differentiate between dialog and narration or different characters. In this work, we propose dialog modeling methods for audiobook synthesis. The proposed approach consists of two stages. First, a text-based dialog style classifier is employed to predict narration vs. dialog from text, and further predict the corresponding characters into soprano and baritone. Then, a dialog style adaptor is added to the text-to-speech (TTS) model to allow synthesizing speech with the corresponding styles. With a speaker verification (SV) based style adaptor, we can even control the strength of a given style. We evaluated the proposed approach in audiobook synthesis with a mean opinion score (MOS) listening test using 9 carefully designed questions. The results show an improvement of 0.35 MOS on dialog distinction without degradation in other aspects. Also a comparative MOS (CMOS) test is conducted to verify the effectiveness of the proposed method.

BibTeX
@inproceedings{icassp2024_dialogmodelingin,
  title = {Dialog Modeling in Audiobook Synthesis},
  author = {Cheng-chieh Yeh and Amirreza Shirani and Weicheng Zhang and Tuomo Raitio and Ramya Rasipuram and Ladan Golipour and David Winarsky},
  booktitle = {ICASSP 2024},
  year = {2024}
}