← Search

Takuma Okamoto

15 accepted papers

2026

WAVENEXT 2: CONVNEXT-BASED FAST NEURAL VOCODERS WITH RESIDUAL DENOISING AND SUB-MODELING FOR GAN AND DIFFUSION MODELS

ICASSP 2026poster

Most neural vocoders are limited to one type: either GAN or diffusion-based. While state-of-the-art models like Vocos and WaveNeXt use powerful ConvNeXt-based generators, they have only been used in GAN frameworks and have limited performance in multi-speaker settings. Moreover, diffusion models, de…

Cited by 0SourcePDFScholar
2025

Mora-Level Prosody Prediction for Text-to-Speech Using Japanese BERT Without Accentual Labels

ICASSP 2025accepted

In practical text-to-speech (TTS) for pitch accent languages, such as Japanese, high-fidelity synthesis with correct prosody requires not only a phoneme sequence but also accentual information. Although accentual information can be obtained from accent dictionaries, words not included in the diction…

Cited by 0SourceScholar
2024

Convnext-TTS And Convnext-VC: Convnext-Based Fast End-To-End Sequence-To-Sequence Text-To-Speech And Voice Conversion

ICASSP 2024accepted

End-to-end (E2E) sequence-to-sequence (S2S) neural text-to-speech (TTS) models and E2E-S2S neural voice conversion (VC) models can achieve high-quality speech synthesis with a single neural network. To further improve the synthesis quality of E2E-S2S TTS and VC models and increase their inference sp…

Cited by 0SourceScholar
2024

FIRNet: Fundamental Frequency Controllable Fast Neural Vocoder With Trainable Finite Impulse Response Filter

ICASSP 2024accepted

Some neural vocoders with fundamental frequency (f <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</inf> ) control have succeeded in performing real-time inference on a single CPU while preserving the quality of the synthetic speech. However, compared…

Cited by 0SourceScholar
2023

Continuous Action Space-Based Spoken Language Acquisition Agent Using Residual Sentence Embedding and Transformer Decoder

ICASSP 2023accepted

Studies on spoken language acquisition agents aim to understand the mechanism of human language learning and to realize it on computers. Existing open vocabulary agents first perform unsupervised word learning from speech signals to construct a word dictionary as a discrete action space and then con…

Cited by 0SourceScholar
2021

High-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VC

ICASSP 2021accepted

This paper presents a high-intelligibility speech synthesis method for persons with dysarthria caused by athetoid cerebral palsy. The muscular control of such speakers is unstable because of their athetoid symptoms, and their pronunciation is unclear, which makes it difficult for them to communicate…

Cited by 0SourceScholar
2021

Noise Level Limited Sub-Modeling for Diffusion Probabilistic Vocoders

ICASSP 2021accepted

Although diffusion probabilistic vocoders WaveGrad and DiffWave can realize real-time high-fidelity speech synthesis with a simple loss function in training, all noise components with over the full range of noise levels are predicted by one model in all iterations. This paper proposes a simple but e…

Cited by 11SourceScholar
2020

Transformer-Based Text-to-Speech with Weighted Forced Attention

ICASSP 2020accepted

This paper investigates state-of-the-art Transformer- and FastSpeech-based high-fidelity neural text-to-speech (TTS) with full-context label input for pitch accent languages. The aim is to realize faster training than conventional Tacotron-based models. Introducing phoneme durations into Tacotron-ba…

Cited by 0SourceScholar
2019

Investigations of Real-time Gaussian Fftnet and Parallel Wavenet Neural Vocoders with Simple Acoustic Features

ICASSP 2019accepted

This paper examines four approaches to improving real-time neural vocoders with simple acoustic features (SAF) constructed from fundamental frequency and mel-cepstra rather than mel-spectrograms. The investigations are as follows: 1) the effectiveness of single Gaussian (SG) autoregressive (AR) Wave…

Cited by 0SourceScholar
2018

An Investigation of Subband Wavenet Vocoder Covering Entire Audible Frequency Range with Limited Acoustic Features

ICASSP 2018accepted

Although a WaveNet vocoder can synthesize more natural-sounding speech waveforms than conventional vocoders with sampling frequencies of 16 and 24 kHz, it is difficult to directly extend the sampling frequency to 48 kHz to cover the entire human audible frequency range for higher-quality synthesis b…

Cited by 0SourceScholar
2017

Analytical approach to 2.5D sound field control using a circular double-layer array of fixed-directivity loudspeakers

ICASSP 2017accepted

This paper provides a regularization-free analytical approach to 2.5D interior and exterior sound field control using a circular double-layer array of fixed-directivity loudspeakers not only to provide a desired sound field inside the array but also to reduce the sound energy outside of it in the ho…

Cited by 0SourceScholar