← Search

Shinsuke Sakai

3 accepted papers

2019

Multi-speaker Sequence-to-sequence Speech Synthesis for Data Augmentation in Acoustic-to-word Speech Recognition

ICASSP 2019accepted

The acoustic-to-word (A2W) automatic speech recognition (ASR) realizes very fast decoding with a simple architecture and achieves state-of-the-art performance. However, the A2W model suffers from the out-of-vocabulary (OOV) word problem and cannot use text-only data to improve the language modeling…

Cited by 0SourceScholar
2017

Semi-supervised ensemble DNN acoustic model training

ICASSP 2017accepted

It is very important to exploit abundant unlabeled speech for improving the acoustic model training in automatic speech recognition (ASR). Semi-supervised training methods incorporate unlabeled data in addition to labeled data to enhance the model training, but it encounters the error-prone label pr…

Cited by 0SourceScholar
2015

Deep autoencoders augmented with phone-class feature for reverberant speech recognition

ICASSP 2015accepted

This paper addresses reverberant speech recognition based on front-end processing using DAE (Deep AutoEncoder) coupled with DNN (Deep Neural Network) acoustic model. DAE can effectively and flexibly learn mapping from corrupted speech to the original clean speech based on the deep learning scheme. W…

Cited by 0SourceScholar