← Search

Kazuhiro Kobayashi

9 accepted papers

2024

Electrolaryngeal Speech Intelligibility Enhancement through Robust Linguistic Encoders

ICASSP 2024accepted

We propose a novel framework for electrolaryngeal speech intelligibility enhancement through the use of robust linguistic encoders. Pretraining and fine-tuning approaches have proven to work well in this task, but in most cases, various mismatches, such as the speech type mismatch (electrolaryngeal…

Cited by 0SourceScholar
2023

Low-Latency Electrolaryngeal Speech Enhancement Based on Fastspeech2-Based Voice Conversion and Self-Supervised Speech Representation

ICASSP 2023accepted

In this paper, we propose a low-latency sequence-to-sequence speech enhancement technique for electrolaryngeal (EL) speech. A low-latency EL speech enhancement technique based on CLDNN was previously proposed to enable laryngectomees to produce relatively naturally sounding speech compared to the or…

Cited by 0SourceScholar
2022

An Investigation of Streaming Non-Autoregressive sequence-to-sequence Voice Conversion

ICASSP 2022accepted

Recent advances in sequence-to-sequence (S2S) models have improved the quality of voice conversion (VC), but it requires the entire sequence to perform inference, which prevents using it in real-time applications. To address this issue, this paper extends the non-autoregressive (NAR) S2S-VC model to…

Cited by 0SourceScholar
2021

Crank: An Open-Source Software for Nonparallel Voice Conversion Based on Vector-Quantized Variational Autoencoder

ICASSP 2021accepted

In this paper, we present an open-source software for developing a nonparallel voice conversion (VC) system named crank. Although we have released an open-source VC software based on the Gaussian mixture model named sprocket in the last VC Challenge, it is not straightforward to apply any speech cor…

Cited by 0SourceScholar
2021

Non-Autoregressive Sequence-To-Sequence Voice Conversion

ICASSP 2021accepted

This paper proposes a novel voice conversion (VC) method based on non-autoregressive sequence-to-sequence (NAR-S2S) models. Inspired by the great success of NAR-S2S models such as FastSpeech in text-to-speech (TTS), we extend the FastSpeech2 model for the VC problem. We introduce the convolution-aug…

Cited by 0SourceScholar
2020

Efficient Shallow Wavenet Vocoder Using Multiple Samples Output Based on Laplacian Distribution and Linear Prediction

ICASSP 2020accepted

This paper presents a novel way for an efficient implementation scheme of shallow WaveNet vocoder with multiple samples (segment) output based on the use of Laplacian distribution and linear prediction. In our previous work, we have proposed a shallow architecture for WaveNet vocoder that utilizes o…

Cited by 0SourceScholar
2019

Voice Conversion with Cyclic Recurrent Neural Network and Fine-tuned Wavenet Vocoder

ICASSP 2019accepted

This paper presents a novel framework for providing high-quality parallel voice conversion (VC) using a cyclic recurrent neural network (RNN) and a finely tuned WaveNet vocoder. Using the proposed system, we are tackling the quality degradation issue faced by WaveNet when it is fed with estimated (o…

Cited by 0SourceScholar
2016

An estimation method of voice timbre evaluation values using feature extraction with Gaussian mixture model based on reference singer

ICASSP 2016accepted

This paper presents an estimation method of voice timbre evaluation values for arbitrary singer's singing voices generated with a singing voice synthesis system towards the development of a singing voice retrieval system. The voice timbre evaluation values are numerical values corresponding to voice…

Cited by 0SourceScholar
2016

Implementation of F0 transformation for statistical singing voice conversion based on direct waveform modification

ICASSP 2016accepted

This paper presents a technique for transforming F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> in a framework of statistical singing voice conversion with direct waveform modification based on spectrum differential (DIFFSVC). The DIFFSVC met…

Cited by 0SourceScholar