← Search

Yamato Ohtani

3 accepted papers

2025

Mora-Level Prosody Prediction for Text-to-Speech Using Japanese BERT Without Accentual Labels

ICASSP 2025accepted

In practical text-to-speech (TTS) for pitch accent languages, such as Japanese, high-fidelity synthesis with correct prosody requires not only a phoneme sequence but also accentual information. Although accentual information can be obtained from accent dictionaries, words not included in the diction…

Cited by 0SourceScholar
2024

Convnext-TTS And Convnext-VC: Convnext-Based Fast End-To-End Sequence-To-Sequence Text-To-Speech And Voice Conversion

ICASSP 2024accepted

End-to-end (E2E) sequence-to-sequence (S2S) neural text-to-speech (TTS) models and E2E-S2S neural voice conversion (VC) models can achieve high-quality speech synthesis with a single neural network. To further improve the synthesis quality of E2E-S2S TTS and VC models and increase their inference sp…

Cited by 0SourceScholar
2024

FIRNet: Fundamental Frequency Controllable Fast Neural Vocoder With Trainable Finite Impulse Response Filter

ICASSP 2024accepted

Some neural vocoders with fundamental frequency (f <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</inf> ) control have succeeded in performing real-time inference on a single CPU while preserving the quality of the synthetic speech. However, compared…

Cited by 0SourceScholar