Precisely Controllable Neural Speech Synthesis
Paul Konstantin Krug, Christoph Wagner, Peter Birkholz, Timo Stich
Abstract
Recent advances in deep learning have significantly improved the quality of speech synthesis, yet these models often suffer from limited controllability and lack of interpretability due to their black-box nature. In contrast, articulatory speech synthesis offers fine-grained control and transparency by simulating sound production based on vocal tract geometry, though it struggles with naturalness and synthesis quality. To bridge these gaps, we propose a novel white-box approach that leverages synthetic articulatory trajectories for neural synthesis, ensuring a fully disentangled, interpretable, and controllable yet high-quality speech synthesis process. Utilizing the VocalTractLab articulatory synthesizer, our method allows the quality of its speech representation to be verified through physical simulation. The proposed system achieves state-of-the-art results in both articulatory and deep articulatory synthesis. To the best of our knowledge, this is the first work to synthesize highly intelligible speech from a purely synthetic articulatory latent representation.
BibTeX
@inproceedings{icassp2025_preciselycontrol,
title = {Precisely Controllable Neural Speech Synthesis},
author = {Paul Konstantin Krug and Christoph Wagner and Peter Birkholz and Timo Stich},
booktitle = {ICASSP 2025},
year = {2025}
}