High-Acoustic Fidelity Text To Speech Synthesis With Fine-Grained Control Of Speech Attributes
Recently developed neural-based TTS models have focused on robustness and finer control over acoustic features such as phoneme duration, energy, and F<inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</inf>, allowing users to have some degree of control…