← Search

Meng Cai

6 accepted papers

2022

Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection

ICASSP 2022accepted

Nowadays, most methods for end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on phrase-level contextual modeling and attention-based relevance modeling, they may suffer from the confusion between simil…

Cited by 59SourceScholar
2022

Improving Pseudo-Label Training For End-To-End Speech Recognition Using Gradient Mask

ICASSP 2022accepted

In the recent trend of semi-supervised speech recognition, both self-supervised representation learning and pseudo-labeling have shown promising results. In this paper, we propose a novel approach to combine their ideas for end-to-end speech recognition model. Without any extra loss function, we uti…

Cited by 0SourceScholar
2021

Improving RNN Transducer Modeling for Small-Footprint Keyword Spotting

ICASSP 2021accepted

The recurrent neural network transducer (RNN-T) model has been proved effective for keyword spotting (KWS) recently. However, compared with cross-entropy (CE) or connectionist temporal classification (CTC) based models, the additional prediction network in the RNN-T model increases the model size an…

Cited by 0SourceScholar
2017

Deep neural networks based speaker modeling at different levels of phonetic granularity

ICASSP 2017accepted

Recently, a hybrid deep neural network/i-vector framework has been proved effective for speaker verification, where the DNN trained to predict tied-triphone states (senones) is used to produce frame alignments for sufficient statistics extraction. In this work, in order to better understand the impa…

Cited by 0SourceScholar
2015

Neuron sparseness versus connection sparseness in deep neural network for large vocabulary speech recognition

ICASSP 2015accepted

Exploiting sparseness in deep neural networks is an important method for reducing the computational cost. In this paper, we study neuron sparseness in deep neural networks for acoustic modeling. For the feed-forward stage, we only activate neurons whose input values are larger than a given threshold…

Cited by 0SourceScholar
2015

The THUEE system for the openKWS14 keyword search evaluation

ICASSP 2015accepted

The OpenKWS14 keyword search evaluation is one of the most challenging and influential evaluations in the field of speech recognition. Its goal is to build a high-performance keyword search system for a minority language with limited training data in a short period of time. We present the system of…

Cited by 0SourceScholar