← Search

Sei Ueno

3 accepted papers

2022

Phone-Informed Refinement of Synthesized Mel Spectrogram for Data Augmentation in Speech Recognition

ICASSP 2022accepted

While recent end-to-end automatic speech recognition (ASR) models achieve high performance, we need to prepare an abundant amount of training data, which is a barrier to apply them to a specific domain. To mitigate the lack of training data, text-to-speech (TTS) systems have been utilized to leverag…

Cited by 0SourceScholar
2019

Multi-speaker Sequence-to-sequence Speech Synthesis for Data Augmentation in Acoustic-to-word Speech Recognition

ICASSP 2019accepted

The acoustic-to-word (A2W) automatic speech recognition (ASR) realizes very fast decoding with a simple architecture and achieves state-of-the-art performance. However, the A2W model suffers from the out-of-vocabulary (OOV) word problem and cannot use text-only data to improve the language modeling…

Cited by 0SourceScholar
2018

Acoustic-to-Word Attention-Based Model Complemented with Character-Level CTC-Based Model

ICASSP 2018accepted

This paper addresses end-to-end speech recognition which directly maps acoustic features to a word sequence. The acoustic-to-word model is attractive since it does not require an external language model and an elaborate decoder, resulting in extremely simple and fast decoding. The apparent drawback…

Cited by 0SourceScholar