← Search

Michael L. Seltzer

18 accepted papers

2024

End-to-End Speech Recognition Contextualization with Large Language Models

ICASSP 2024accepted

In recent years, Large Language Models (LLMs) have garnered significant attention from the research community due to their exceptional performance and generalization capabilities. In this paper, we introduce a novel method for contextualizing speech recognition models incorporating LLMs. Our approac…

Cited by 0SourceScholar
2023

Factorized Blank Thresholding for Improved Runtime Efficiency of Neural Transducers

ICASSP 2023accepted

We show how factoring the RNN-T’s output distribution can significantly reduce the computation cost and power consumption for on-device ASR inference with no loss in accuracy. With the rise in popularity of neural-transducer type models like the RNN-T for on-device ASR, optimizing RNN-T’s runtime ef…

Cited by 0SourceScholar
2023

Improving fast-slow Encoder based Transducer with Streaming Deliberation

ICASSP 2023accepted

This paper introduces a fast-slow encoder based transducer with streaming deliberation for end-to-end automatic speech recognition. We aim to improve the recognition accuracy of the fast-slow encoder based transducer while keeping its latency low by integrating a streaming deliberation model. Specif…

Cited by 0SourceScholar
2023

Massively Multilingual ASR on 70 Languages: Tokenization, Architecture, and Generalization Capabilities

ICASSP 2023accepted

End-to-end multilingual ASR has become more appealing because of several reasons such as simplifying the training and deployment process and positive performance transfer from high-resource to low-resource languages. However, scaling up the number of languages, total hours, and number of unique toke…

Cited by 0SourceScholar
2022

Neural-FST Class Language Model for End-to-End Speech Recognition

ICASSP 2022accepted

We propose Neural-FST Class Language Model (NFCLM) for end-to-end speech recognition, a novel method that combines neural network language models (NNLMs) and finite state transducers (FSTs) in a mathematically consistent framework. Our method utilizes a background NNLM which models generic backgroun…

Cited by 0SourceScholar
2021

Improved Neural Language Model Fusion for Streaming Recurrent Neural Network Transducer

ICASSP 2021accepted

Recurrent Neural Network Transducer (RNN-T), like most end-to-end speech recognition model architectures, has an implicit neural network language model (NNLM) and cannot easily leverage unpaired text data during training. Previous work has proposed various fusion methods to incorporate external NNLM…

Cited by 0SourceScholar
2021

Memory-Efficient Speech Recognition on Smart Devices

ICASSP 2021accepted

Recurrent transducer models have emerged as a promising solution for speech recognition on the current and next generation smart devices. The transducer models provide competitive accuracy within a reasonable memory footprint alleviating the memory capacity constraints in these devices. However, the…

Cited by 0SourceScholar
2020

Aipnet: Generative Adversarial Pre-Training of Accent-Invariant Networks for End-To-End Speech Recognition

ICASSP 2020accepted

As one of the major sources in speech variability, accents have posed a grand challenge to the robustness of speech recognition systems. In this paper, our goal is to build a unified end-to-end speech recognition system that generalizes well across accents. For this purpose, we propose a novel pre-t…

Cited by 0SourceScholar
2020

G2G: TTS-Driven Pronunciation Learning for Graphemic Hybrid ASR

ICASSP 2020accepted

Grapheme-based acoustic modeling has recently been shown to outperform phoneme-based approaches in both hybrid and end-to-end automatic speech recognition (ASR), even on non-phonemic languages like English. However, graphemic ASR still has problems with low-frequency words that do not follow the sta…

Cited by 0SourceScholar
2020

Transformer-Based Acoustic Modeling for Hybrid Speech Recognition

ICASSP 2020accepted

We propose and evaluate transformer-based acoustic models (AMs) for hybrid speech recognition. Several modeling choices are discussed in this work, including various positional embedding methods and an iterated loss to enable training deep transformers. We also present a preliminary study of using l…

Cited by 0SourceScholar
2019

End-to-end Contextual Speech Recognition Using Class Language Models and a Token Passing Decoder

ICASSP 2019accepted

End-to-end modeling (E2E) of automatic speech recognition (ASR) blends all the components of a traditional speech recognition system into a single, unified model. Although it simplifies the ASR systems, the unified model is hard to adapt when training and testing data mismatches. In this work, we fo…

Cited by 0SourceScholar
2018

Efficient Integration of Fixed Beamformers and Speech Separation Networks for Multi-Channel Far-Field Speech Separation

ICASSP 2018accepted

Speech separation research has significantly progressed in recent years thanks to the rapid advances in deep learning technology. However the performance of recently proposed single-channel neural network-based speech separation methods is still limited especially in reverberant environments. To pus…

Cited by 0SourceScholar
2017

A study on data augmentation of reverberant speech for robust speech recognition

ICASSP 2017accepted

The environmental robustness of DNN-based acoustic models can be significantly improved by using multi-condition training data. However, as data collection is a costly proposition, simulation of the desired conditions is a frequently adopted strategy. In this paper we detail a data augmentation appr…

Cited by 0SourceScholar
2016

Deep beamforming networks for multi-channel speech recognition

ICASSP 2016accepted

Despite the significant progress in speech recognition enabled by deep neural networks, poor performance persists in some scenarios. In this work, we focus on far-field speech recognition which remains challenging due to high levels of noise and reverberation in the captured speech signals. We propo…

Cited by 0SourceScholar
2015

Improving speech recognition in reverberation using a room-aware deep neural network and multi-task learning

ICASSP 2015accepted

In this paper, we propose two approaches to improve deep neural network (DNN) acoustic models for speech recognition in reverberant environments. Both methods utilize auxiliary information in training the DNN but differ in the type of information and the manner in which it is used. The first method…

Cited by 0SourceScholar
2015

Speech recognition with prediction-adaptation-correction recurrent neural networks

ICASSP 2015accepted

We propose the prediction-adaptation-correction RNN (PAC-RNN), in which a correction DNN estimates the state posterior probability based on both the current frame and the prediction made on the past frames by a prediction DNN. The result from the main DNN is fed back to the prediction DNN to make be…

Cited by 0SourceScholar