← Search

Yusuke Ijima

11 accepted papers

2024

Noise-Robust Zero-Shot Text-to-Speech Synthesis Conditioned on Self-Supervised Speech-Representation Model with Adapters

ICASSP 2024accepted

The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately. However, this approach suffers from degradation in speech synthesis quality when…

Cited by 0SourceScholar
2024

STYLECAP: Automatic Speaking-Style Captioning from Speech Based on Speech and Language Self-Supervised Learning Models

ICASSP 2024accepted

We propose StyleCap, a method to generate natural language descriptions of speaking styles appearing in speech. Although most of conventional techniques for para-/non-linguistic information recognition focus on the category classification or the intensity estimation of pre-defined labels, they canno…

Cited by 0SourceScholar
2024

What Do Self-Supervised Speech and Speaker Models Learn? New Findings from a Cross Model Layer-Wise Analysis

ICASSP 2024accepted

Self-supervised learning (SSL) has attracted increased attention for learning meaningful speech representations. Speech SSL models, such as WavLM, employ masked prediction training to encode general-purpose representations. In contrast, speaker SSL models, exemplified by DINO-based models, adopt utt…

Cited by 0SourceScholar
2023

Enhancement of Text-Predicting Style Token With Generative Adversarial Network for Expressive Speech Synthesis

ICASSP 2023accepted

This work proposes an advanced text-predicting style embedding for expressive speech synthesis. Text-predicting global style token (TPGST) predicts style embedding from text instead of reference speech and uses it to condition a text-to-speech synthesis (TTS) model, resulting in style TTS without re…

Cited by 0SourceScholar
2021

Simpleflat: A Simple Whole-Network Pre-Training Approach for RNN Transducer-Based End-to-End Speech Recognition

ICASSP 2021accepted

Recurrent neural network-transducer (RNN-T) is promising for building time-synchronous end-to-end automatic speech recognition (ASR) systems, in part because it does not need frame-wise alignment between input features and target labels in the training step. Although training without alignment is be…

Cited by 8SourceScholar
2021

Speech Emotion Recognition Based on Listener Adaptive Models

ICASSP 2021accepted

This paper presents a novel speech emotion recognition scheme that can deal with the individuality of emotion perception. Most conventional methods directly model the majority decision of multiple listener’s perceived emotions. However, emotion perception varies with the listener, which means the co…

Cited by 0SourceScholar
2018

Neural Confnet Classification: Fully Neural Network Based Spoken Utterance Classification Using Word Confusion Networks

ICASSP 2018accepted

This paper describes neural ConfNet classification, a novel fully neural network based spoken utterance classification method that uses word confusion networks (ConfNets). Our motivation is to establish a spoken utterance classification method that can precisely understand natural language and robus…

Cited by 0SourceScholar
2018

Non-Parallel Voice Conversion Using Variational Autoencoders Conditioned by Phonetic Posteriorgrams and D-Vectors

ICASSP 2018accepted

This paper proposes novel frameworks for non-parallel voice conversion (VC) using variational autoencoders (VAEs). Although conventional VAE-based VC models can be trained using non-parallel speech corpora with given speaker representations, phonetic contents of the converted speech tend to vanish b…

Cited by 114SourceScholar
2018

Soft-Target Training with Ambiguous Emotional Utterances for DNN-Based Speech Emotion Classification

ICASSP 2018accepted

This paper presents a novel emotion classification method for natural speech. One of the problems in the state-of-the-art method based on Deep Neural Network (DNN) is the paucity of the training data compared to model complexity. To solve this problem, this paper utilizes the ambiguous emotional utt…

Cited by 0SourceScholar
2017

Generative adversarial network-based postfilter for statistical parametric speech synthesis

ICASSP 2017accepted

We propose a postfilter based on a generative adversarial network (GAN) to compensate for the differences between natural speech and speech synthesized by statistical parametric speech synthesis. In particular, we focus on the differences caused by over-smoothing, which makes the sounds muffled. Ove…

Cited by 0SourceScholar