← Search

Hisashi Kawai

26 accepted papers

2026

WAVENEXT 2: CONVNEXT-BASED FAST NEURAL VOCODERS WITH RESIDUAL DENOISING AND SUB-MODELING FOR GAN AND DIFFUSION MODELS

ICASSP 2026poster

Most neural vocoders are limited to one type: either GAN or diffusion-based. While state-of-the-art models like Vocos and WaveNeXt use powerful ConvNeXt-based generators, they have only been used in GAN frameworks and have limited performance in multi-speaker settings. Moreover, diffusion models, de…

Cited by 0SourcePDFScholar
2025

Mora-Level Prosody Prediction for Text-to-Speech Using Japanese BERT Without Accentual Labels

ICASSP 2025accepted

In practical text-to-speech (TTS) for pitch accent languages, such as Japanese, high-fidelity synthesis with correct prosody requires not only a phoneme sequence but also accentual information. Although accentual information can be obtained from accent dictionaries, words not included in the diction…

Cited by 0SourceScholar
2024

Convnext-TTS And Convnext-VC: Convnext-Based Fast End-To-End Sequence-To-Sequence Text-To-Speech And Voice Conversion

ICASSP 2024accepted

End-to-end (E2E) sequence-to-sequence (S2S) neural text-to-speech (TTS) models and E2E-S2S neural voice conversion (VC) models can achieve high-quality speech synthesis with a single neural network. To further improve the synthesis quality of E2E-S2S TTS and VC models and increase their inference sp…

Cited by 0SourceScholar
2024

FIRNet: Fundamental Frequency Controllable Fast Neural Vocoder With Trainable Finite Impulse Response Filter

ICASSP 2024accepted

Some neural vocoders with fundamental frequency (f <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</inf> ) control have succeeded in performing real-time inference on a single CPU while preserving the quality of the synthetic speech. However, compared…

Cited by 0SourceScholar
2024

Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-Based ASR

ICASSP 2024accepted

Due to the modality discrepancy between textual and acoustic modeling, efficiently transferring linguistic knowledge from a pretrained language model (PLM) to acoustic encoding for automatic speech recognition (ASR) still remains a challenging task. In this study, we propose a cross-modality knowled…

Cited by 0SourceScholar
2021

CrossMap Transformer: A Crossmodal Masked Path Transformer Using Double Back-Translation for Vision-and-Language Navigation

RA-L 2021

Navigation guided by natural language instructions is particularly suitable for Domestic Service Robots that interacts naturally with users. This task involves the prediction of a sequence of actions that leads to a specified destination given a natural language navigation instruction. The task thus

Cited by 15SourceScholar
2021

High-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VC

ICASSP 2021accepted

This paper presents a high-intelligibility speech synthesis method for persons with dysarthria caused by athetoid cerebral palsy. The muscular control of such speakers is unstable because of their athetoid symptoms, and their pronunciation is unclear, which makes it difficult for them to communicate…

Cited by 0SourceScholar
2021

Noise Level Limited Sub-Modeling for Diffusion Probabilistic Vocoders

ICASSP 2021accepted

Although diffusion probabilistic vocoders WaveGrad and DiffWave can realize real-time high-fidelity speech synthesis with a simple loss function in training, all noise components with over the full range of noise levels are predicted by one model in all iterations. This paper proposes a simple but e…

Cited by 0SourceScholar
2021

Unsupervised Neural Adaptation Model Based on Optimal Transport for Spoken Language Identification

ICASSP 2021accepted

Due to the mismatch of statistical distributions of acoustic speech between training and testing sets, the performance of spoken language identification (SLID) could be drastically degraded. In this paper, we propose an unsupervised neural adaptation model to deal with the distribution mismatch prob…

Cited by 0SourceScholar
2020

A Multimodal Target-Source Classifier With Attention Branches to Understand Ambiguous Instructions for Fetching Daily Objects

RA-L 2020

In this study, we focus on multimodal language understanding for fetching instructions in the domestic service robots context. This task consists of predicting a target object, as instructed by the user, given an image and an unstructured sentence, such as “Bring me the yellow box (from the wooden c

Cited by 10SourceScholar
2020

Alleviating the Burden of Labeling: Sentence Generation by Attention Branch Encoder-Decoder Network

RA-L 2020

Domestic service robots (DSRs) are a promising solution to the shortage of home care workers. However, one of the main limitations of DSRs is their inability to interact naturally through language. Recently, data-driven approaches have been shown to be effective for tackling this limitation; however

Cited by 12SourceScholar
2020

Transformer-Based Text-to-Speech with Weighted Forced Attention

ICASSP 2020accepted

This paper investigates state-of-the-art Transformer- and FastSpeech-based high-fidelity neural text-to-speech (TTS) with full-context label input for pitch accent languages. The aim is to realize faster training than conventional Tacotron-based models. Introducing phoneme durations into Tacotron-ba…

Cited by 0SourceScholar
2019

Interactive Learning of Teacher-student Model for Short Utterance Spoken Language Identification

ICASSP 2019accepted

Short utterance-based spoken language identification (LID) is a challenging task due to the large variation of its feature representation. Improving feature representation of short utterances using a teacher-student method has been shown its effectiveness for LID tasks. However, conventional teacher…

Cited by 0SourceScholar
2019

Investigation of Sequence-level Knowledge Distillation Methods for CTC Acoustic Models

ICASSP 2019accepted

This paper presents knowledge distillation (KD) methods for training connectionist temporal classification (CTC) acoustic models. In a previous study, we proposed a KD method based on the sequence-level cross-entropy, and showed that the conventional KD method based on the frame-level cross-entropy…

Cited by 0SourceScholar
2019

Investigations of Real-time Gaussian Fftnet and Parallel Wavenet Neural Vocoders with Simple Acoustic Features

ICASSP 2019accepted

This paper examines four approaches to improving real-time neural vocoders with simple acoustic features (SAF) constructed from fundamental frequency and mel-cepstra rather than mel-spectrograms. The investigations are as follows: 1) the effectiveness of single Gaussian (SG) autoregressive (AR) Wave…

Cited by 0SourceScholar
2019

Multimodal Attention Branch Network for Perspective-Free Sentence Generation

CoRL 2019

In this paper, we address the automatic sentence generation of fetching instructions for domestic service robots. Typical fetching commands such as “bring me the yellow toy from the upper part of the white shelf” includes referring expressions, i.e., “from the white upper part of the white shelf”. T

Cited by 0SourcePDFScholar
2019

Understanding Natural Language Instructions for Fetching Daily Objects Using GAN-Based Multimodal Target-Source Classification

RA-L 2019

In this letter, we address multimodal language understanding with unconstrained fetching instruction for domestic service robots. A typical fetching instruction such as “Bring me the yellow toy from the white shelf” requires to infer the user intention, i.e., what object (target) to fetch and from w

Cited by 35SourceScholar
2018

A Multimodal Classifier Generative Adversarial Network for Carry and Place Tasks From Ambiguous Language Instructions

RA-L 2018

This letter focuses on a multimodal language understanding method for carry-and-place tasks with domestic service robots. We address the case of ambiguous instructions, that is, when the target area is not specified. For instance “put away the milk and cereal” is a natural instruction where there is

Cited by 31SourceScholar
2018

An Investigation of Noise Shaping with Perceptual Weighting for Wavenet-Based Speech Generation

ICASSP 2018accepted

We propose a noise shaping method to improve the sound quality of speech signals generated by WaveNet, which is a convolutional neural network (CNN) that predicts a waveform sample sequence as a discrete symbol sequence. Speech signals generated by WaveNet often suffer from noise signals caused by t…

Cited by 0SourceScholar
2018

An Investigation of Subband Wavenet Vocoder Covering Entire Audible Frequency Range with Limited Acoustic Features

ICASSP 2018accepted

Although a WaveNet vocoder can synthesize more natural-sounding speech waveforms than conventional vocoders with sampling frequencies of 16 and 24 kHz, it is difficult to directly extend the sampling frequency to 48 kHz to cover the entire human audible frequency range for higher-quality synthesis b…

Cited by 0SourceScholar
2018

An Investigation of a Knowledge Distillation Method for CTC Acoustic Models

ICASSP 2018accepted

End-to-end acoustic models, such as connectionist temporal classification (CTC) and the attention model, have been studied, and their speech recognition accuracies come close to those of conventional deep neural network (DNN)-hidden Markov models. However, most high-performance end-to-end models are…

Cited by 0SourceScholar
2018

Comparative Evaluations of Various Factored Deep Convolutional Rnn Architectures for Noise Robust Speech Recognition

ICASSP 2018accepted

In this paper, we present a factored network-based acoustic modeling framework with various deep convolutional recurrent neural network (RNN) architectures for noise-robust automatic speech recognition (ASR). As the factored network-based acoustic model, we have already proposed a deep convolutional…

Cited by 0SourceScholar
2017

Minimum Bayes risk training of CTC acoustic models in maximum a posteriori based decoding framework

ICASSP 2017accepted

When using connectionist temporal classification (CTC) based acoustic models (AMs) for large vocabulary continuous speech recognition (LVCSR), most previous studies have used a naive interpolation of the CTC-AM score and an additional language model score, although there is no theoretical justificat…

Cited by 0SourceScholar
2016

Bottleneck linear transformation network adaptation for speaker adaptive training-based hybrid DNN-HMM speech recognizer

ICASSP 2016accepted

Recently, a Hybrid DNN-HMM recognizer trained with the Speaker Adaptive Training (SAT) concept was successfully modified to a more effective speaker-adaptation-oriented recognizer whose DNN front-end adopted a Linear Transformation Network (LTN) Speaker Dependent (SD) module. However, the size of SD…

Cited by 0SourceScholar
2016

Local fisher discriminant analysis for spoken language identification

ICASSP 2016accepted

I-vector is a state-of-the-art technique widely used in spoken language identification systems. Since i-vectors include total variability factors, discriminant analysis methods have been introduced to find the most discriminative features while removing the undesired variables for language identific…

Cited by 0SourceScholar