← Search

Masato Mimura

12 accepted papers

2025

Advancing Streaming ASR with Chunk-wise Attention and Trans-chunk Selective State Spaces

ICASSP 2025accepted

This paper explores enhancing streaming speech recognition through the integration of chunk-wise attention and selective state space models (SSMs). The proposed framework replaces the quadratic complexity of attention-based context incorporation with a fully recurrent module based on selective SSMs.…

Cited by 0SourceScholar
2025

Alignment-Free Training for Transducer-based Multi-Talker ASR

ICASSP 2025accepted

Extending the RNN Transducer (RNNT) to recognize multi-talker speech is essential for wider automatic speech recognition (ASR) applications. Multi-talker RNNT (MT-RNNT) aims to achieve recognition without relying on costly front-end source separation. MT-RNNT is conventionally implemented using arch…

Cited by 0SourceScholar
2025

Leveraging IPA and Articulatory Features as Effective Inductive Biases for Multilingual ASR Training

ICASSP 2025accepted

In recent advancements in end-to-end ASR, large-scale self-supervised or weakly supervised models have achieved a significant milestone. However, it remains challenging to train consistently high-performing multilingual models, transferable to languages without much resource. In this study, we propo…

Cited by 0SourceScholar
2023

Time-Domain Speech Enhancement Assisted by Multi-Resolution Frequency Encoder and Decoder

ICASSP 2023accepted

Time-domain speech enhancement (SE) has recently been intensively investigated. Among recent works, DEMUCS [1] introduces multi-resolution STFT loss to enhance performance. However, some resolutions used for STFT contain non-stationary signals, and it is challenging to learn multi-resolution frequen…

Cited by 0SourceScholar
2022

Selective Multi-Task Learning For Speech Emotion Recognition Using Corpora Of Different Styles

ICASSP 2022accepted

While speech emotion recognition (SER) has been actively studied, the amount and variations of training data are limited compared with speech recognition and speaker recognition tasks. Therefore, it is promising to combine multiple corpora to train a generalized SER model. However, the manner of emo…

Cited by 0SourceScholar
2019

Multi-speaker Sequence-to-sequence Speech Synthesis for Data Augmentation in Acoustic-to-word Speech Recognition

ICASSP 2019accepted

The acoustic-to-word (A2W) automatic speech recognition (ASR) realizes very fast decoding with a simple architecture and achieves state-of-the-art performance. However, the A2W model suffers from the out-of-vocabulary (OOV) word problem and cannot use text-only data to improve the language modeling…

Cited by 0SourceScholar
2018

Acoustic-to-Word Attention-Based Model Complemented with Character-Level CTC-Based Model

ICASSP 2018accepted

This paper addresses end-to-end speech recognition which directly maps acoustic features to a word sequence. The acoustic-to-word model is attractive since it does not require an external language model and an elaborate decoder, resulting in extremely simple and fast decoding. The apparent drawback…

Cited by 0SourceScholar
2018

An End-to-End Approach to Joint Social Signal Detection and Automatic Speech Recognition

ICASSP 2018accepted

Social signals such as laughter and fillers are often observed in natural conversation, and they play various roles in human-to-human communication. Detecting these events is useful for transcription systems to generate rich transcription and for dialogue systems to behave as we do such as synchroni…

Cited by 0SourceScholar
2018

Statistical Speech Enhancement Based on Probabilistic Integration of Variational Autoencoder and Non-Negative Matrix Factorization

ICASSP 2018accepted

This paper presents a statistical method of single-channel speech enhancement that uses a variational autoencoder (VAE) as a prior distribution on clean speech. A standard approach to speech enhancement is to train a deep neural network (DNN) to take noisy speech as input and output clean speech. Al…

Cited by 0SourceScholar
2018

Unsupervised Beamforming Based on Multichannel Nonnegative Matrix Factorization for Noisy Speech Recognition

ICASSP 2018accepted

This paper presents unsupervised multichannel speech enhancement for noisy speech recognition. Time-frequency (TF) mask estimation has actively been studied for estimating the steering vectors and spatial covariance matrices of speech and noise used for beamforming. The state-of-the-art approach to…

Cited by 0SourceScholar
2017

Semi-supervised ensemble DNN acoustic model training

ICASSP 2017accepted

It is very important to exploit abundant unlabeled speech for improving the acoustic model training in automatic speech recognition (ASR). Semi-supervised training methods incorporate unlabeled data in addition to labeled data to enhance the model training, but it encounters the error-prone label pr…

Cited by 0SourceScholar
2015

Deep autoencoders augmented with phone-class feature for reverberant speech recognition

ICASSP 2015accepted

This paper addresses reverberant speech recognition based on front-end processing using DAE (Deep AutoEncoder) coupled with DNN (Deep Neural Network) acoustic model. DAE can effectively and flexibly learn mapping from corrupted speech to the original clean speech based on the deep learning scheme. W…

Cited by 0SourceScholar