← Search

Shigeki Karita

6 accepted papers

2022

Knowledge Transfer from Large-Scale Pretrained Language Models to End-To-End Speech Recognizers

ICASSP 2022accepted

End-to-end speech recognition is a promising technology for enabling compact automatic speech recognition (ASR) systems since it can unify the acoustic and language model into a single neural network. However, as a drawback, training of end-to-end speech recognizers always requires transcribed utter…

Cited by 0SourceScholar
2019

Semi-supervised End-to-end Speech Recognition Using Text-to-speech and Autoencoders

ICASSP 2019accepted

We introduce speech and text autoencoders that share encoders and decoders with an automatic speech recognition (ASR) model to improve ASR performance with large speech only and text only training datasets. To build the speech and text autoencoders, we leverage state-of-the-art ASR and text-to-speec…

Cited by 0SourceScholar
2018

Frame-by-Frame Closed-Form Update for Mask-Based Adaptive MVDR Beamforming

ICASSP 2018accepted

Beamforming approaches using time-frequency masks have recently been investigated and have shown promising results for noise robust automatic speech recognition (ASR) in many tasks. The time-frequency masks are estimated to compute the spatial statistics of target speech and noise signals, and then…

Cited by 0SourceScholar
2018

Rescoring N-Best Speech Recognition List Based on One-on-One Hypothesis Comparison Using Encoder-Classifier Model

ICASSP 2018accepted

This paper proposes a new model for accurately rescoring (reranking) N-best speech recognition hypothesis lists. The model is based on state-of-the-art neural networks (NNs) and provides the minimum necessary functionality to perform N-best rescoring, i.e. one-on-one hypothesis comparison on a given…

Cited by 0SourceScholar
2018

Sequence Training of Encoder-Decoder Model Using Policy Gradient for End-to-End Speech Recognition

ICASSP 2018accepted

The standard evaluation metric of automatic speech recognition (ASR) is the word error rate (WER), which measures the dissimilarity between recognized word sequences and their ground truth. Many training algorithms designed to reduce sequence-level errors such as WER have been proposed for hidden Ma…

Cited by 0SourceScholar
2015

Far-field speech recognition using CNN-DNN-HMM with convolution in time

ICASSP 2015accepted

Recent studies in speech recognition have shown that the performance of convolutional neural networks (CNNs) is superior to that of fully connected deep neural networks (DNNs). In this paper, we explore the use of CNNs in far-field speech recognition for dealing with reverberation, which blurs spect…

Cited by 0SourceScholar