← Search

Ryuichi Yamamoto

15 accepted papers

2025

Description-Based Controllable Text-to-Speech With Cross-Lingual Voice Control

ICASSP 2025accepted

We propose a novel description-based controllable text-to-speech (TTS) method with cross-lingual control capability. To address the lack of audio-description paired data in the target language, we combine a TTS model trained on the target language with a description control model trained on another…

Cited by 0SourceScholar
2025

Investigating Factors Related to the Naturalness of Synthesized Unison Singing

ICASSP 2025accepted

Singing voice synthesis (SVS) technology has progressed rapidly in recent years. However, vocal ensemble synthesis has not yet been widely explored. In this work, we focus on unison singing, which is to have several singers singing the same melody together. Our goal is to understand what acoustic pr…

Cited by 0SourceScholar
2024

Electrolaryngeal Speech Intelligibility Enhancement through Robust Linguistic Encoders

ICASSP 2024accepted

We propose a novel framework for electrolaryngeal speech intelligibility enhancement through the use of robust linguistic encoders. Pretraining and fine-tuning approaches have proven to work well in this task, but in most cases, various mismatches, such as the speech type mismatch (electrolaryngeal…

Cited by 0SourceScholar
2024

Enhancing Multilingual TTS with Voice Conversion Based Data Augmentation and Posterior Embedding

ICASSP 2024accepted

This paper proposes a multilingual, multi-speaker (MM) TTS system by using a voice conversion (VC)-based data augmentation method. Creating an MM-TTS model is challenging, owing to the difficulties of collecting polyglot data from multiple speakers. To address this problem, we adopt a cross-lingual,…

Cited by 0SourceScholar
2024

PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-To-Speech Using Natural Language Descriptions

ICASSP 2024accepted

We propose PromptTTS++, a prompt-based text-to-speech (TTS) synthesis system that allows control over speaker identity using natural language descriptions. To control speaker identity within the prompt-based TTS framework, we introduce the concept of speaker prompt, which describes voice characteris…

Cited by 0SourceScholar
2023

Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform

ICASSP 2023accepted

We propose a lightweight end-to-end text-to-speech model using multi-band generation and inverse short-time Fourier transform. Our model is based on VITS, a high-quality end-to-end text-to-speech model, but adopts two changes for more efficient inference: 1) the most computationally expensive compon…

Cited by 0SourceScholar
2023

Nonparallel High-Quality Audio Super Resolution with Domain Adaptation and Resampling CycleGANs

ICASSP 2023accepted

Neural audio super-resolution models are typically trained on low- and high-resolution audio signal pairs. Although these methods achieve highly accurate super-resolution if the acoustic characteristics of the input data are similar to those of the training data, challenges remain: the models suffer…

Cited by 0SourceScholar
2023

Period VITS: Variational Inference with Explicit Pitch Modeling for End-To-End Emotional Speech Synthesis

ICASSP 2023accepted

Several fully end-to-end text-to-speech (TTS) models have been proposed that have shown better performance compared to cascade models (i.e., training acoustic and vocoder models separately). However, they often generate unstable pitch contour with audible artifacts when the dataset contains emotiona…

Cited by 0SourceScholar
2021

Parallel Waveform Synthesis Based on Generative Adversarial Networks with Voicing-Aware Conditional Discriminators

ICASSP 2021accepted

This paper proposes voicing-aware conditional discriminators for Parallel WaveGAN-based waveform synthesis systems. In this framework, we adopt a projection-based conditioning method that can significantly improve the discriminator’s performance. Furthermore, the conventional discriminator is separa…

Cited by 19SourceScholar
2021

TTS-by-TTS: TTS-Driven Data Augmentation for Fast and High-Quality Speech Synthesis

ICASSP 2021accepted

In this paper, we propose a text-to-speech (TTS)-driven data augmentation method for improving the quality of a non-autoregressive (AR) TTS system. Recently proposed non-AR models, such as FastSpeech 2, have successfully achieved fast speech synthesis system. However, their quality is not satisfacto…

Cited by 0SourceScholar
2020

Espnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit

ICASSP 2020accepted

This paper introduces a new end-to-end text-to-speech (E2E-TTS) toolkit named ESPnet-TTS, which is an extension of the open-source speech processing toolkit ESPnet. The toolkit supports state-of- the-art E2E-TTS models, including Tacotron 2, Transformer TTS, and FastSpeech, and also provides recipes…

Cited by 0SourceScholar
2020

Improving LPCNET-Based Text-to-Speech with Linear Prediction-Structured Mixture Density Network

ICASSP 2020accepted

In this paper, we propose an improved LPCNet vocoder using a linear prediction (LP)-structured mixture density network (MDN). The recently proposed LPCNet vocoder has successfully achieved high-quality and lightweight speech synthesis systems by combining a vocal tract LP filter with a WaveRNN-based…

Cited by 0SourceScholar
2020

Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram

ICASSP 2020accepted

We propose Parallel WaveGAN, a distillation-free, fast, and small-footprint waveform generation method using a generative adversarial network. In the proposed method, a non-autoregressive WaveNet is trained by jointly optimizing multi-resolution spectrogram and adversarial loss functions, which can…

Cited by 1020SourceScholar
2020

Semi-Supervised Speaker Adaptation for End-to-End Speech Synthesis with Pretrained Models

ICASSP 2020accepted

Recently, end-to-end text-to-speech (TTS) models have achieved a remarkable performance, however, requiring a large amount of paired text and speech data for training. On the other hand, we can easily collect unpaired dozen minutes of speech recordings for a target speaker without corresponding text…

Cited by 0SourceScholar