← Search

Ken'ichi Kumatani

8 accepted papers

2021

Ensemble Combination between Different Time Segmentations

ICASSP 2021accepted

Hypothesis-level combination between multiple models can often yield gains in speech recognition. However, all models in the ensemble are usually restricted to use the same audio segmentation times. This paper proposes to generalise hypothesis-level combination, allowing the use of different audio s…

Cited by 0SourceScholar
2020

Fully Learnable Front-End for Multi-Channel Acoustic Modeling Using Semi-Supervised Learning

ICASSP 2020accepted

In this work, we investigated the teacher-student training paradigm to train a fully learnable multi-channel acoustic model for far-field automatic speech recognition (ASR). Using a large offline teacher model trained on beamformed audio, we trained a simpler multi-channel student acoustic model use…

Cited by 0SourceScholar
2020

Robust Multi-Channel Speech Recognition Using Frequency Aligned Network

ICASSP 2020accepted

Conventional speech enhancement technique such as beamforming has known benefits for far-field speech recognition. Our own work in frequency-domain multi-channel acoustic modeling has shown additional improvements by training a spatial filtering layer jointly within an acoustic model. In this paper,…

Cited by 0SourceScholar
2019

Frequency Domain Multi-channel Acoustic Modeling for Distant Speech Recognition

ICASSP 2019accepted

Conventional far-field automatic speech recognition (ASR) systems typically employ microphone array techniques for speech enhancement in order to improve robustness against noise or reverberation. However, such speech enhancement techniques do not always yield ASR accuracy improvement because the op…

Cited by 39SourceScholar
2019

Improving Noise Robustness of Automatic Speech Recognition via Parallel Data and Teacher-student Learning

ICASSP 2019accepted

For real-world speech recognition applications, noise robustness is still a challenge. In this work, we adopt the teacher-student (T/S) learning technique using a parallel clean and noisy corpus for improving automatic speech recognition (ASR) performance under multimedia noise. On top of that, we a…

Cited by 53SourceScholar
2019

Multi-geometry Spatial Acoustic Modeling for Distant Speech Recognition

ICASSP 2019accepted

The use of spatial information with multiple microphones can improve far-field automatic speech recognition (ASR) accuracy. However, conventional microphone array techniques degrade speech enhancement performance when there is an array geometry mismatch between design and test conditions. Moreover,…

Cited by 18SourceScholar
2018

Time-Delayed Bottleneck Highway Networks Using a DFT Feature for Keyword Spotting

ICASSP 2018accepted

This paper presents a novel deep neural network (DNN) architecture with highway blocks (HWs) using a complex discrete Fourier transform (DFT) feature for keyword spotting. In our previous work, we showed that the feed-forward DNN with a time-delayed bottleneck layer (TDB-DNN) directly trained from t…

Cited by 0SourceScholar