← Search

Kentaro Tachibana

7 accepted papers

2025

Description-Based Controllable Text-to-Speech With Cross-Lingual Voice Control

ICASSP 2025accepted

We propose a novel description-based controllable text-to-speech (TTS) method with cross-lingual control capability. To address the lack of audio-description paired data in the target language, we combine a TTS model trained on the target language with a description control model trained on another…

Cited by 0SourceScholar
2024

PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-To-Speech Using Natural Language Descriptions

ICASSP 2024accepted

We propose PromptTTS++, a prompt-based text-to-speech (TTS) synthesis system that allows control over speaker identity using natural language descriptions. To control speaker identity within the prompt-based TTS framework, we introduce the concept of speaker prompt, which describes voice characteris…

Cited by 0SourceScholar
2023

Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform

ICASSP 2023accepted

We propose a lightweight end-to-end text-to-speech model using multi-band generation and inverse short-time Fourier transform. Our model is based on VITS, a high-quality end-to-end text-to-speech model, but adopts two changes for more efficient inference: 1) the most computationally expensive compon…

Cited by 0SourceScholar
2023

Nonparallel High-Quality Audio Super Resolution with Domain Adaptation and Resampling CycleGANs

ICASSP 2023accepted

Neural audio super-resolution models are typically trained on low- and high-resolution audio signal pairs. Although these methods achieve highly accurate super-resolution if the acoustic characteristics of the input data are similar to those of the training data, challenges remain: the models suffer…

Cited by 0SourceScholar
2023

Period VITS: Variational Inference with Explicit Pitch Modeling for End-To-End Emotional Speech Synthesis

ICASSP 2023accepted

Several fully end-to-end text-to-speech (TTS) models have been proposed that have shown better performance compared to cascade models (i.e., training acoustic and vocoder models separately). However, they often generate unstable pitch contour with audible artifacts when the dataset contains emotiona…

Cited by 0SourceScholar
2018

An Investigation of Noise Shaping with Perceptual Weighting for Wavenet-Based Speech Generation

ICASSP 2018accepted

We propose a noise shaping method to improve the sound quality of speech signals generated by WaveNet, which is a convolutional neural network (CNN) that predicts a waveform sample sequence as a discrete symbol sequence. Speech signals generated by WaveNet often suffer from noise signals caused by t…

Cited by 0SourceScholar
2018

An Investigation of Subband Wavenet Vocoder Covering Entire Audible Frequency Range with Limited Acoustic Features

ICASSP 2018accepted

Although a WaveNet vocoder can synthesize more natural-sounding speech waveforms than conventional vocoders with sampling frequencies of 16 and 24 kHz, it is difficult to directly extend the sampling frequency to 48 kHz to cover the entire human audible frequency range for higher-quality synthesis b…

Cited by 0SourceScholar