Hierarchical Prosody Modeling and Control in Non-Autoregressive Parallel Neural TTS
Neural text-to-speech (TTS) synthesis can generate speech that is indistinguishable from natural speech. However, the synthetic speech often represents the average prosodic style of the database instead of having more versatile prosodic variation. Moreover, many models lack the ability to control th…