← Search

Hirokazu Kameoka

31 accepted papers

2025

Rethinking Mean Opinion Scores in Speech Quality Assessment: Score Aggregation through Quantized Distribution Fitting

ICASSP 2025accepted

This study addresses the task of speech quality assessment (SQA), which aims to automatically predict the subjective quality of a given speech. Recent efforts have focused on training neural-based models to predict the mean opinion score (MOS) of speech samples produced by text-to-speech or voice co…

Cited by 0SourceScholar
2024

Selecting N-Lowest Scores for Training MOS Prediction Models

ICASSP 2024accepted

The automatic speech quality assessment (SQA) has been extensively studied to predict the speech quality without time-consuming questionnaires. Recently, neural-based SQA models have been actively developed for speech samples produced by text-to-speech or voice conversion, with a primary focus on tr…

Cited by 0SourceScholar
2024

Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator

ICASSP 2024accepted

A generative adversarial network (GAN)-based vocoder trained with an adversarial discriminator is commonly used for speech synthesis because of its fast, lightweight, and high-quality characteristics. However, this data-driven model requires a large amount of training data incurring high data-collec…

Cited by 0SourceScholar
2023

JSV-VC: Jointly Trained Speaker Verification and Voice Conversion Models

ICASSP 2023accepted

This paper proposes a variational autoencoder (VAE)-based method for voice conversion (VC) on arbitrary source-target speaker pairs without parallel corpora, i.e., non-parallel any-to-any VC. One typical approach is to use speaker embeddings obtained from a speaker verification (SV) model as the con…

Cited by 0SourceScholar
2023

Wave-U-Net Discriminator: Fast and Lightweight Discriminator for Generative Adversarial Network-Based Speech Synthesis

ICASSP 2023accepted

In speech synthesis, a generative adversarial network (GAN), training a generator (speech synthesizer) and a discriminator in a min-max game, is widely used to improve speech quality. An ensemble of discriminators is commonly used in recent neural vocoders (e.g., HiFi-GAN) and end-to-end text-to-spe…

Cited by 0SourceScholar
2022

Attentionpit: Soft Permutation Invariant Training for Audio Source Separation with Attention Mechanism

ICASSP 2022accepted

Permutation invariant training (PIT) has recently attracted attention as a framework to achieve end-to-end time-domain audio source separation. Its goal is to train a separation network that takes a mixture signal as input and produces the J underlying source signals. Since the order of the output s…

Cited by 0SourceScholar
2022

HBP: An Efficient Block Permutation Solver Using Hungarian Algorithm and Spectrogram Inpainting for Multichannel Audio Source Separation

ICASSP 2022accepted

This paper proposes a method called "Hungarian Block Permutation (HBP)" to solve the block permutation problem in frequency-domain multichannel audio source separation. Many methods for frequency-domain multichannel audio source separation are designed to simultaneously solve frequency-wise source s…

Cited by 0SourceScholar
2022

ISTFTNET: Fast and Lightweight Mel-Spectrogram Vocoder Incorporating Inverse Short-Time Fourier Transform

ICASSP 2022accepted

In recent text-to-speech synthesis and voice conversion systems, a mel-spectrogram is commonly applied as an intermediate representation, and the necessity for a mel-spectrogram vocoder is increasing. A mel-spectrogram vocoder must solve three inverse problems: recovery of the original-scale magnitu…

Cited by 0SourceScholar
2022

Investigation And Comparison of Optimization Methods for Variational Autoencoder-Based Underdetermined Multichannel Source Separation

ICASSP 2022accepted

In this paper, we investigate two algorithms for variational autoencoder (VAE)-based underdetermined multichannel source separation. We previously extended the multichannel VAE (MVAE) method for determined multichannel source separation and proposed the generalized MVAE (GMVAE) method for underdeter…

Cited by 0SourceScholar
2021

Maskcyclegan-VC: Learning Non-Parallel Voice Conversion with Filling in Frames

ICASSP 2021accepted

Non-parallel voice conversion (VC) is a technique for training voice converters without a parallel corpus. Cycle-consistent adversarial network-based VCs (CycleGAN-VC and CycleGAN-VC2) are widely accepted as benchmark methods. However, owing to their insufficient ability to grasp time-frequency stru…

Cited by 0SourceScholar
2021

SepNet: A Deep Separation Matrix Prediction Network for Multichannel Audio Source Separation

ICASSP 2021accepted

In this paper, we propose SepNet, a deep neural network (DNN) designed to predict separation matrices from multichannel observations. One well-known approach to blind source separation (BSS) involves independent component analysis (ICA). A recently developed method called independent low-rank matrix…

Cited by 0SourceScholar
2019

ATTS2S-VC: Sequence-to-sequence Voice Conversion with Attention and Context Preservation Mechanisms

ICASSP 2019accepted

This paper describes a method based on a sequence-to-sequence learning (Seq2Seq) with attention and context preservation mechanism for voice conversion (VC) tasks. Seq2Seq has been outstanding at numerous tasks involving sequence modeling such as speech synthesis and recognition, machine translation…

Cited by 0SourceScholar
2019

Cyclegan-VC2: Improved Cyclegan-based Non-parallel Voice Conversion

ICASSP 2019accepted

Non-parallel voice conversion (VC) is a technique for learning the mapping from source to target speech without relying on parallel data. This is an important task, but it has been challenging due to the disadvantages of the training conditions. Recently, CycleGAN-VC has provided a breakthrough and…

Cited by 0SourceScholar
2019

Fast MVAE: Joint Separation and Classification of Mixed Sources Based on Multichannel Variational Autoencoder with Auxiliary Classifier

ICASSP 2019accepted

This paper proposes an alternative algorithm for the multi-channel variational autoencoder (MVAE), a recently proposed multichannel source separation approach. While MVAE is notable for its impressive source separation performance, its convergence-guaranteed optimization algorithm and the fact that…

Cited by 0SourceScholar
2019

Joint Separation and Dereverberation of Reverberant Mixtures with Multichannel Variational Autoencoder

ICASSP 2019accepted

In this paper, we deal with a multichannel source separation problem under a highly reverberant condition. The multichannel variational autoencoder (MVAE) is a recently proposed source separation method that employs the decoder distribution of a conditional VAE (CVAE) as the generative model for the…

Cited by 0SourceScholar
2019

Seeing through Sounds: Predicting Visual Semantic Segmentation Results from Multichannel Audio Signals

ICASSP 2019accepted

Sounds provide us with vast amounts of information about surrounding objects and can even remind us visual images of them. Is it possible to implement this noteworthy human ability on machines? In this paper, we study a new task that consists of predicting image recognition results in the form of se…

Cited by 0SourceScholar
2018

Joint Separation and Dereverberation of Reverberant Mixtures with Determined Multichannel Non-Negative Matrix Factorization

ICASSP 2018accepted

This paper proposes an extension of multichannel non-negative matrix factorization (MNMF) that simultaneously solves source separation and dereverberation. While MNMF was originally formulated under an underdetermined problem setting where sources can outnumber microphones, a determined counterpart…

Cited by 0SourceScholar
2018

Speech Waveform Synthesis from MFCC Sequences with Generative Adversarial Networks

ICASSP 2018accepted

This paper proposes a method for generating speech from filterbank mel frequency cepstral coefficients (MFCC), which are widely used in speech applications, such as ASR, but are generally considered unusable for speech synthesis. First, we predict fundamental frequency and voicing information from M…

Cited by 0SourceScholar
2018

Vae-Space: Deep Generative Model of Voice Fundamental Frequency Contours

ICASSP 2018accepted

Modeling the speech generation process can provide flexible and interpretable ways to generate intended synthetic speech. In this paper, we present a deep generative model of fundamental frequency (F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</su…

Cited by 7SourceScholar
2017

A majorization-minimization algorithm with projected gradient updates for time-domain spectrogram factorization

ICASSP 2017accepted

We previously introduced a framework called time-domain spectrogram factorization (TSF), which realizes nonnegative matrix factorization (NMF)-like source separation in the time domain. This framework is particularly noteworthy in that, while maintaining the ability of NMF to obtain a parts-based re…

Cited by 0SourceScholar
2017

A noise suppression method for body-conducted soft speech based on non-negative tensor factorization of air- and body-conducted signals

ICASSP 2017accepted

This paper presents a novel noise suppression method to enhance soft speech recorded with a special body-conductive microphone called nonaudible murmur (NAM) microphone. NAM microphone is capable of detecting extremely soft speech, but the recorded soft speech easily suffers from external noise due…

Cited by 0SourceScholar
2017

Fast algorithm for statistical phrase/accent command estimation based on generative model incorporating spectral features

ICASSP 2017accepted

An important challenge in speech processing involves extracting non-linguistic information from a fundamental frequency (F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> ) contour of speech. We propose a fast algorithm for estimating the model…

Cited by 0SourceScholar
2017

Generative adversarial network-based postfilter for statistical parametric speech synthesis

ICASSP 2017accepted

We propose a postfilter based on a generative adversarial network (GAN) to compensate for the differences between natural speech and speech synthesized by statistical parametric speech synthesis. In particular, we focus on the differences caused by over-smoothing, which makes the sounds muffled. Ove…

Cited by 0SourceScholar
2016

Shifted and convolutive source-filter non-negative matrix factorization for monaural audio source separation

ICASSP 2016accepted

This paper proposes an extension of non-negative matrix factorization (NMF), which combines the shifted NMF model with the source-filter model. Shifted NMF was proposed as a powerful approach for monaural source separation and multiple fundamental frequency (F0) estimation, which is particularly uni…

Cited by 0SourceScholar
2016

Sparse sound field decomposition with multichannel extension of complex NMF

ICASSP 2016accepted

A sparse sound field decomposition method using prior information on source signals in the time-frequency domain is proposed. Sparse sound field decomposition has been proved to be effective for various acoustic signal processing applications. Current methods for sparse decomposition are based only…

Cited by 0SourceScholar
2016

Statistical F0 prediction for electrolaryngeal speech enhancement considering generative process of F0 contours within product of experts framework

ICASSP 2016accepted

We have previously proposed a statistical fundamental frequency (F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> ) prediction method that makes it possible to predict the underlying F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:x…

Cited by 0SourceScholar
2015

Efficient multichannel nonnegative matrix factorization exploiting rank-1 spatial model

ICASSP 2015accepted

This paper proposes a new efficient multichannel nonnegative matrix factorization (NMF) method. Recently, multichannel NMF (MNMF) has been proposed as a means of solving the blind source separation problem. This method estimates a mixing system of sources and attempts to separate them in a blind fas…

Cited by 0SourceScholar
2015

Lp-norm non-negative matrix factorization and its application to singing voice enhancement

ICASSP 2015accepted

Measures of sparsity are useful in many aspects of audio signal processing including speech enhancement, audio coding and singing voice enhancement, and the well-known method for these applications is non-negative matrix factorization (NMF), which decomposes a non-negative data matrix into two non-n…

Cited by 0SourceScholar