← Search

Kou Tanaka

11 accepted papers

2025

Rethinking Mean Opinion Scores in Speech Quality Assessment: Score Aggregation through Quantized Distribution Fitting

ICASSP 2025accepted

This study addresses the task of speech quality assessment (SQA), which aims to automatically predict the subjective quality of a given speech. Recent efforts have focused on training neural-based models to predict the mean opinion score (MOS) of speech samples produced by text-to-speech or voice co…

Cited by 0SourceScholar
2024

Selecting N-Lowest Scores for Training MOS Prediction Models

ICASSP 2024accepted

The automatic speech quality assessment (SQA) has been extensively studied to predict the speech quality without time-consuming questionnaires. Recently, neural-based SQA models have been actively developed for speech samples produced by text-to-speech or voice conversion, with a primary focus on tr…

Cited by 0SourceScholar
2024

Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator

ICASSP 2024accepted

A generative adversarial network (GAN)-based vocoder trained with an adversarial discriminator is commonly used for speech synthesis because of its fast, lightweight, and high-quality characteristics. However, this data-driven model requires a large amount of training data incurring high data-collec…

Cited by 0SourceScholar
2023

JSV-VC: Jointly Trained Speaker Verification and Voice Conversion Models

ICASSP 2023accepted

This paper proposes a variational autoencoder (VAE)-based method for voice conversion (VC) on arbitrary source-target speaker pairs without parallel corpora, i.e., non-parallel any-to-any VC. One typical approach is to use speaker embeddings obtained from a speaker verification (SV) model as the con…

Cited by 0SourceScholar
2023

Wave-U-Net Discriminator: Fast and Lightweight Discriminator for Generative Adversarial Network-Based Speech Synthesis

ICASSP 2023accepted

In speech synthesis, a generative adversarial network (GAN), training a generator (speech synthesizer) and a discriminator in a min-max game, is widely used to improve speech quality. An ensemble of discriminators is commonly used in recent neural vocoders (e.g., HiFi-GAN) and end-to-end text-to-spe…

Cited by 12SourceScholar
2022

ISTFTNET: Fast and Lightweight Mel-Spectrogram Vocoder Incorporating Inverse Short-Time Fourier Transform

ICASSP 2022accepted

In recent text-to-speech synthesis and voice conversion systems, a mel-spectrogram is commonly applied as an intermediate representation, and the necessity for a mel-spectrogram vocoder is increasing. A mel-spectrogram vocoder must solve three inverse problems: recovery of the original-scale magnitu…

Cited by 0SourceScholar
2021

Maskcyclegan-VC: Learning Non-Parallel Voice Conversion with Filling in Frames

ICASSP 2021accepted

Non-parallel voice conversion (VC) is a technique for training voice converters without a parallel corpus. Cycle-consistent adversarial network-based VCs (CycleGAN-VC and CycleGAN-VC2) are widely accepted as benchmark methods. However, owing to their insufficient ability to grasp time-frequency stru…

Cited by 0SourceScholar
2019

ATTS2S-VC: Sequence-to-sequence Voice Conversion with Attention and Context Preservation Mechanisms

ICASSP 2019accepted

This paper describes a method based on a sequence-to-sequence learning (Seq2Seq) with attention and context preservation mechanism for voice conversion (VC) tasks. Seq2Seq has been outstanding at numerous tasks involving sequence modeling such as speech synthesis and recognition, machine translation…

Cited by 115SourceScholar
2019

Cyclegan-VC2: Improved Cyclegan-based Non-parallel Voice Conversion

ICASSP 2019accepted

Non-parallel voice conversion (VC) is a technique for learning the mapping from source to target speech without relying on parallel data. This is an important task, but it has been challenging due to the disadvantages of the training conditions. Recently, CycleGAN-VC has provided a breakthrough and…

Cited by 0SourceScholar
2018

Vae-Space: Deep Generative Model of Voice Fundamental Frequency Contours

ICASSP 2018accepted

Modeling the speech generation process can provide flexible and interpretable ways to generate intended synthetic speech. In this paper, we present a deep generative model of fundamental frequency (F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</su…

Cited by 7SourceScholar
2016

Statistical F0 prediction for electrolaryngeal speech enhancement considering generative process of F0 contours within product of experts framework

ICASSP 2016accepted

We have previously proposed a statistical fundamental frequency (F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> ) prediction method that makes it possible to predict the underlying F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:x…

Cited by 0SourceScholar