← Search

Tomoki Toda

48 accepted papers

2025

Improvements of Discriminative Feature Space Training for Anomalous Sound Detection in Unlabeled Conditions

ICASSP 2025accepted

In anomalous sound detection, the discriminative method has demonstrated superior performance. This approach constructs a discriminative feature space through the classification of the meta-information labels for normal sounds. This feature space reflects the differences in machine sounds and effect…

Cited by 0SourceScholar
2025

Investigating Factors Related to the Naturalness of Synthesized Unison Singing

ICASSP 2025accepted

Singing voice synthesis (SVS) technology has progressed rapidly in recent years. However, vocal ensemble synthesis has not yet been widely explored. In this work, we focus on unison singing, which is to have several singers singing the same melody together. Our goal is to understand what acoustic pr…

Cited by 0SourceScholar
2025

Mora-Level Prosody Prediction for Text-to-Speech Using Japanese BERT Without Accentual Labels

ICASSP 2025accepted

In practical text-to-speech (TTS) for pitch accent languages, such as Japanese, high-fidelity synthesis with correct prosody requires not only a phoneme sequence but also accentual information. Although accentual information can be obtained from accent dictionaries, words not included in the diction…

Cited by 0SourceScholar
2024

Convnext-TTS And Convnext-VC: Convnext-Based Fast End-To-End Sequence-To-Sequence Text-To-Speech And Voice Conversion

ICASSP 2024accepted

End-to-end (E2E) sequence-to-sequence (S2S) neural text-to-speech (TTS) models and E2E-S2S neural voice conversion (VC) models can achieve high-quality speech synthesis with a single neural network. To further improve the synthesis quality of E2E-S2S TTS and VC models and increase their inference sp…

Cited by 0SourceScholar
2024

Electrolaryngeal Speech Intelligibility Enhancement through Robust Linguistic Encoders

ICASSP 2024accepted

We propose a novel framework for electrolaryngeal speech intelligibility enhancement through the use of robust linguistic encoders. Pretraining and fine-tuning approaches have proven to work well in this task, but in most cases, various mismatches, such as the speech type mismatch (electrolaryngeal…

Cited by 0SourceScholar
2024

FIRNet: Fundamental Frequency Controllable Fast Neural Vocoder With Trainable Finite Impulse Response Filter

ICASSP 2024accepted

Some neural vocoders with fundamental frequency (f <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</inf> ) control have succeeded in performing real-time inference on a single CPU while preserving the quality of the synthetic speech. However, compared…

Cited by 0SourceScholar
2024

MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction

ICASSP 2024accepted

The prevalent approach in speech emotion recognition (SER) involves integrating both audio and textual information to comprehensively identify the speaker’s emotion, with the text generally obtained through automatic speech recognition (ASR). An essential issue of this approach is that ASR errors fr…

Cited by 0SourceScholar
2023

Intermediate Fine-Tuning Using Imperfect Synthetic Speech for Improving Electrolaryngeal Speech Recognition

ICASSP 2023accepted

Research on automatic speech recognition (ASR) systems for electrolaryngeal speakers has been relatively unexplored due to small datasets. When training data is lacking in ASR, a large-scale pre-training and fine tuning framework is often sufficient to achieve high recognition rates; however, in ele…

Cited by 0SourceScholar
2023

Low-Latency Electrolaryngeal Speech Enhancement Based on Fastspeech2-Based Voice Conversion and Self-Supervised Speech Representation

ICASSP 2023accepted

In this paper, we propose a low-latency sequence-to-sequence speech enhancement technique for electrolaryngeal (EL) speech. A low-latency EL speech enhancement technique based on CLDNN was previously proposed to enable laryngectomees to produce relatively naturally sounding speech compared to the or…

Cited by 0SourceScholar
2023

Source-Filter HiFi-GAN: Fast and Pitch Controllable High-Fidelity Neural Vocoder

ICASSP 2023accepted

Our previous work, the unified source-filter GAN (uSFGAN) vocoder, introduced a novel architecture based on the source- filter theory into the parallel waveform generative adversarial network to achieve high voice quality and pitch controllability. However, the high temporal resolution inputs result…

Cited by 0SourceScholar
2023

Text-To-Speech Synthesis Based on Latent Variable Conversion Using Diffusion Probabilistic Model and Variational Autoencoder

ICASSP 2023accepted

Text-to-speech synthesis (TTS) is a task to convert texts into speech. Two of the factors that have been driving TTS are the advancements of probabilistic models and latent representation learning. We propose a TTS method based on latent variable conversion using a diffusion probabilistic model and…

Cited by 0SourceScholar
2022

An Investigation of Streaming Non-Autoregressive sequence-to-sequence Voice Conversion

ICASSP 2022accepted

Recent advances in sequence-to-sequence (S2S) models have improved the quality of voice conversion (VC), but it requires the entire sequence to perform inference, which prevents using it in real-time applications. To address this issue, this paper extends the non-autoregressive (NAR) S2S-VC model to…

Cited by 0SourceScholar
2022

Direct Noisy Speech Modeling for Noisy-To-Noisy Voice Conversion

ICASSP 2022accepted

Beyond the conventional voice conversion (VC) where the speaker information is converted without altering the linguistic content, the background sounds are informative and need to be retained in some real-world scenarios, such as VC in movie/video and VC in music where the voice is entangled with ba…

Cited by 0SourceScholar
2022

LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech

ICASSP 2022accepted

An effective approach to automatically predict the subjective rating for synthetic speech is to train on a listening test dataset with human-annotated scores. Although each speech sample in the dataset is rated by several listeners, most previous works only used the mean score as the training target…

Cited by 0SourceScholar
2022

S3PRL-VC: Open-Source Voice Conversion Framework with Self-Supervised Speech Representations

ICASSP 2022accepted

This paper introduces S3PRL-VC, an open-source voice conversion (VC) framework based on the S3PRL toolkit. In the context of recognition-synthesis VC, self-supervised speech representation (S3R) is valuable in its potential to replace the expensive supervised representation adopted by state-of-the-a…

Cited by 0SourceScholar
2022

Towards Identity Preserving Normal to Dysarthric Voice Conversion

ICASSP 2022accepted

We present a voice conversion framework that converts normal speech into dysarthric speech while preserving the speaker identity. Such a framework is essential for (1) clinical decision making processes and alleviation of patient stress, (2) data augmentation for dysarthric speech recognition. This…

Cited by 0SourceScholar
2021

Crank: An Open-Source Software for Nonparallel Voice Conversion Based on Vector-Quantized Variational Autoencoder

ICASSP 2021accepted

In this paper, we present an open-source software for developing a nonparallel voice conversion (VC) system named crank. Although we have released an open-source VC software based on the Gaussian mixture model named sprocket in the last VC Challenge, it is not straightforward to apply any speech cor…

Cited by 0SourceScholar
2021

High-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VC

ICASSP 2021accepted

This paper presents a high-intelligibility speech synthesis method for persons with dysarthria caused by athetoid cerebral palsy. The muscular control of such speakers is unstable because of their athetoid symptoms, and their pronunciation is unclear, which makes it difficult for them to communicate…

Cited by 0SourceScholar
2021

Noise Level Limited Sub-Modeling for Diffusion Probabilistic Vocoders

ICASSP 2021accepted

Although diffusion probabilistic vocoders WaveGrad and DiffWave can realize real-time high-fidelity speech synthesis with a simple loss function in training, all noise components with over the full range of noise levels are predicted by one model in all iterations. This paper proposes a simple but e…

Cited by 0SourceScholar
2021

Non-Autoregressive Sequence-To-Sequence Voice Conversion

ICASSP 2021accepted

This paper proposes a novel voice conversion (VC) method based on non-autoregressive sequence-to-sequence (NAR-S2S) models. Inspired by the great success of NAR-S2S models such as FastSpeech in text-to-speech (TTS), we extend the FastSpeech2 model for the VC problem. We introduce the convolution-aug…

Cited by 0SourceScholar
2021

Speech Emotion Recognition Based on Listener Adaptive Models

ICASSP 2021accepted

This paper presents a novel speech emotion recognition scheme that can deal with the individuality of emotion perception. Most conventional methods directly model the majority decision of multiple listener’s perceived emotions. However, emotion perception varies with the listener, which means the co…

Cited by 0SourceScholar
2021

Speech Recognition by Simply Fine-Tuning Bert

ICASSP 2021accepted

We propose a simple method for automatic speech recognition (ASR) by fine-tuning BERT, which is a language model (LM) trained on large-scale unlabeled text data and can generate rich contextual representations. Our assumption is that given a history context sequence, a powerful LM can narrow the ran…

Cited by 0SourceScholar
2020

Efficient Shallow Wavenet Vocoder Using Multiple Samples Output Based on Laplacian Distribution and Linear Prediction

ICASSP 2020accepted

This paper presents a novel way for an efficient implementation scheme of shallow WaveNet vocoder with multiple samples (segment) output based on the use of Laplacian distribution and linear prediction. In our previous work, we have proposed a shallow architecture for WaveNet vocoder that utilizes o…

Cited by 0SourceScholar
2020

Espnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit

ICASSP 2020accepted

This paper introduces a new end-to-end text-to-speech (E2E-TTS) toolkit named ESPnet-TTS, which is an extension of the open-source speech processing toolkit ESPnet. The toolkit supports state-of- the-art E2E-TTS models, including Tacotron 2, Transformer TTS, and FastSpeech, and also provides recipes…

Cited by 0SourceScholar
2020

Transformer-Based Text-to-Speech with Weighted Forced Attention

ICASSP 2020accepted

This paper investigates state-of-the-art Transformer- and FastSpeech-based high-fidelity neural text-to-speech (TTS) with full-context label input for pitch accent languages. The aim is to realize faster training than conventional Tacotron-based models. Introducing phoneme durations into Tacotron-ba…

Cited by 0SourceScholar
2020

Weakly-Supervised Sound Event Detection with Self-Attention

ICASSP 2020accepted

In this paper, we propose a novel sound event detection (SED) method that incorporates a self-attention mechanism of the Transformer for a weakly-supervised learning scenario. The proposed method utilizes the Transformer encoder, which consists of multiple self-attention modules, allowing to take bo…

Cited by 0SourceScholar
2019

Investigations of Real-time Gaussian Fftnet and Parallel Wavenet Neural Vocoders with Simple Acoustic Features

ICASSP 2019accepted

This paper examines four approaches to improving real-time neural vocoders with simple acoustic features (SAF) constructed from fundamental frequency and mel-cepstra rather than mel-spectrograms. The investigations are as follows: 1) the effectiveness of single Gaussian (SG) autoregressive (AR) Wave…

Cited by 0SourceScholar
2019

Scene-dependent Anomalous Acoustic-event Detection Based on Conditional Wavenet and I-vector

ICASSP 2019accepted

This paper proposes a scene-dependent anomalous acoustic-event detection based on conditional WaveNet and i-vector. The WaveNet builds normal acoustic event models by exhaustive learning of time-domain signals in the public space to provide scene-independent anomaly detection. I-vectors are used as…

Cited by 0SourceScholar
2019

Voice Conversion with Cyclic Recurrent Neural Network and Fine-tuned Wavenet Vocoder

ICASSP 2019accepted

This paper presents a novel framework for providing high-quality parallel voice conversion (VC) using a cyclic recurrent neural network (RNN) and a finely tuned WaveNet vocoder. Using the proposed system, we are tackling the quality degradation issue faced by WaveNet when it is fed with estimated (o…

Cited by 0SourceScholar
2018

An Investigation of Noise Shaping with Perceptual Weighting for Wavenet-Based Speech Generation

ICASSP 2018accepted

We propose a noise shaping method to improve the sound quality of speech signals generated by WaveNet, which is a convolutional neural network (CNN) that predicts a waveform sample sequence as a discrete symbol sequence. Speech signals generated by WaveNet often suffer from noise signals caused by t…

Cited by 0SourceScholar
2018

An Investigation of Subband Wavenet Vocoder Covering Entire Audible Frequency Range with Limited Acoustic Features

ICASSP 2018accepted

Although a WaveNet vocoder can synthesize more natural-sounding speech waveforms than conventional vocoders with sampling frequencies of 16 and 24 kHz, it is difficult to directly extend the sampling frequency to 48 kHz to cover the entire human audible frequency range for higher-quality synthesis b…

Cited by 0SourceScholar
2017

A noise suppression method for body-conducted soft speech based on non-negative tensor factorization of air- and body-conducted signals

ICASSP 2017accepted

This paper presents a novel noise suppression method to enhance soft speech recorded with a special body-conductive microphone called nonaudible murmur (NAM) microphone. NAM microphone is capable of detecting extremely soft speech, but the recorded soft speech easily suffers from external noise due…

Cited by 0SourceScholar
2017

BLSTM-HMM hybrid system combined with sound activity detection network for polyphonic Sound Event Detection

ICASSP 2017accepted

This paper presents a new hybrid approach for polyphonic Sound Event Detection (SED) which incorporates a temporal structure modeling technique based on a hidden Markov model (HMM) with a frame-by-frame detection method based on a bidirectional long short-term memory (BLSTM) recurrent neural network…

Cited by 0SourceScholar
2016

An estimation method of voice timbre evaluation values using feature extraction with Gaussian mixture model based on reference singer

ICASSP 2016accepted

This paper presents an estimation method of voice timbre evaluation values for arbitrary singer's singing voices generated with a singing voice synthesis system towards the development of a singing voice retrieval system. The voice timbre evaluation values are numerical values corresponding to voice…

Cited by 0SourceScholar
2016

Implementation of F0 transformation for statistical singing voice conversion based on direct waveform modification

ICASSP 2016accepted

This paper presents a technique for transforming F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> in a framework of statistical singing voice conversion with direct waveform modification based on spectrum differential (DIFFSVC). The DIFFSVC met…

Cited by 0SourceScholar
2016

Noise suppression method for body-conducted soft speech enhancement based on external noise monitoring

ICASSP 2016accepted

This paper presents a novel approach to suppressing adverse effects of external noise on body-conducted soft speech for silent speech communication in noisy environments. Nonaudible murmur (NAM) microphone as one of the body-conductive microphones is capable of detecting very soft speech. However, b…

Cited by 0SourceScholar
2016

Statistical F0 prediction for electrolaryngeal speech enhancement considering generative process of F0 contours within product of experts framework

ICASSP 2016accepted

We have previously proposed a statistical fundamental frequency (F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> ) prediction method that makes it possible to predict the underlying F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:x…

Cited by 0SourceScholar
2015

Combination of two-dimensional cochleogram and spectrogram features for deep learning-based ASR

ICASSP 2015accepted

This paper explores the use of auditory features based on cochleograms; two dimensional speech features derived from gammatone filters within the convolutional neural network (CNN) framework. Furthermore, we also propose various possibilities to combine cochleogram features with log-mel filter banks…

Cited by 0SourceScholar
2015

EEG signal enhancement using multi-channel wiener filter with a spatial correlation prior

ICASSP 2015accepted

Event-related potentials (ERPs) of electroencephalogram (EEG) are often used as features for brain machine interfaces or for analysis of brain activities. However, as EEG signals easily suffer from various artifacts, ERPs are often collapsed and hard to observe. There are several attempts at using m…

Cited by 0SourceScholar
2015

Modulation spectrum-constrained trajectory training algorithm for GMM-based Voice Conversion

ICASSP 2015accepted

This paper presents a novel training algorithm for Gaussian Mixture Model (GMM)-based Voice Conversion (VC). One of the advantages of GMM-based VC is computationally efficient conversion processing enabling to achieve real-time VC applications. On the other hand, the quality of the converted speech…

Cited by 0SourceScholar
2015

Parameter generation algorithm considering Modulation Spectrum for HMM-based speech synthesis

ICASSP 2015accepted

This paper proposes a novel parameter generation algorithm for high-quality speech generation in Hidden Markov Model (HMM)-based speech synthesis. One of the biggest issues causing significant quality degradation is the over-smoothing effect often observed in generated parameter trajectories. Global…

Cited by 0SourceScholar
2015

SAS: A speaker verification spoofing database containing diverse attacks

ICASSP 2015accepted

This paper presents the first version of a speaker verification spoofing and anti-spoofing database, named SAS corpus. The corpus includes nine spoofing techniques, two of which are speech synthesis, and seven are voice conversion. We design two protocols, one for standard speaker verification evalu…

Cited by 0SourceScholar