2025
Mora-Level Prosody Prediction for Text-to-Speech Using Japanese BERT Without Accentual Labels
ICASSP 2025accepted
In practical text-to-speech (TTS) for pitch accent languages, such as Japanese, high-fidelity synthesis with correct prosody requires not only a phoneme sequence but also accentual information. Although accentual information can be obtained from accent dictionaries, words not included in the diction…