← Search

Takenori Yoshimura

5 accepted papers

2023

Embedding a Differentiable Mel-Cepstral Synthesis Filter to a Neural Speech Synthesis System

ICASSP 2023accepted

This paper integrates a classic mel-cepstral synthesis filter into a modern neural speech synthesis system towards end-to-end controllable speech synthesis. Since the mel-cepstral synthesis filter is explicitly embedded in neural waveform models in the proposed system, both voice characteristics and…

Cited by 0SourceScholar
2020

End-to-End Automatic Speech Recognition Integrated with CTC-Based Voice Activity Detection

ICASSP 2020accepted

This paper integrates a voice activity detection (VAD) function with end-to-end automatic speech recognition toward an online speech interface and transcribing very long audio recordings. We focus on connectionist temporal classification (CTC) and its extension of CTC/attention architectures. As opp…

Cited by 0SourceScholar
2020

Espnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit

ICASSP 2020accepted

This paper introduces a new end-to-end text-to-speech (E2E-TTS) toolkit named ESPnet-TTS, which is an extension of the open-source speech processing toolkit ESPnet. The toolkit supports state-of- the-art E2E-TTS models, including Tacotron 2, Transformer TTS, and FastSpeech, and also provides recipes…

Cited by 0SourceScholar
2019

Speaker-dependent Wavenet-based Delay-free Adpcm Speech Coding

ICASSP 2019accepted

This paper proposes a WaveNet-based delay-free adaptive differential pulse code modulation (ADPCM) speech coding system. The WaveNet generative model, which is a state-of-the-art model for neural-network-based speech waveform synthesis, is used as the adaptive predictor in ADPCM. To further improve…

Cited by 0SourceScholar
2018

Statistical Voice Conversion Based on Wavenet

ICASSP 2018accepted

This paper proposes a voice conversion technique based on WaveNet to directly generate target audio waveforms from acoustic features of a source speaker. In voice conversion based on statistical models, the relation between acoustic features, such as spectral parameters, extracted from source and ta…

Cited by 0SourceScholar