← Search

Grant P. Strimel

13 accepted papers

2024

Max-Margin Transducer Loss: Improving Sequence-Discriminative Training Using a Large-Margin Learning Strategy

ICASSP 2024accepted

In this work, we propose a novel sequence-discriminative training criterion for automatic speech recognition (ASR) based on the Conformer Transducer. Inspired by the large-margin classifier framework, we separate the "good" and the "bad" hypotheses in an N-best list produced from a pre-trained trans…

Cited by 0SourceScholar
2023

Dialog Act Guided Contextual Adapter for Personalized Speech Recognition

ICASSP 2023accepted

Personalization in multi-turn dialogs has been a long standing challenge for end-to-end automatic speech recognition (E2E ASR) models. Recent work on contextual adapters has tackled rare word recognition using user catalogs. This adaptation, however, does not incorporate an important cue, the dialog…

Cited by 0SourceScholar
2023

Dual-Attention Neural Transducers for Efficient Wake Word Spotting in Speech Recognition

ICASSP 2023accepted

We present dual-attention neural biasing, an architecture designed to boost Wake Words (WW) recognition and improve inference time latency on speech recognition tasks. This architecture enables a dynamic switch for its runtime compute paths by exploiting WW spotting to select which branch of its att…

Cited by 0SourceScholar
2023

Gated Contextual Adapters For Selective Contextual Biasing In Neural Transducers

ICASSP 2023accepted

Neural contextual biasing for end-to-end neural ASR transducers has shown significant improvements in the recognition of named entities, such as contact names or device names. However, it comes with the cost of increased compute, as the biasing layers (which are usually based on cross-attention) add…

Cited by 0SourceScholar
2023

Multilingual End-To-End Spoken Language Understanding For Ultra-Low Footprint Applications

ICASSP 2023accepted

Tiny Signal-to-Interpretation (TinyS2I) has been recently introduced as an ultra low-footprint end-to-end spoken language understanding (SLU) model. This architecture is capable of running in ultra resource constrained environments like voice assistant devices, while at the same time reducing latenc…

Cited by 0SourceScholar
2023

Procter: Pronunciation-Aware Contextual Adapter For Personalized Speech Recognition In Neural Transducers

ICASSP 2023accepted

End-to-End (E2E) automatic speech recognition (ASR) systems used in voice assistants often have difficulties recognizing infrequent words personalized to the user, such as names and places. Rare words often have non-trivial pronunciations, and in such cases, human knowledge in the form of a pronunci…

Cited by 0SourceScholar
2023

Robust Acoustic And Semantic Contextual Biasing In Neural Transducers For Speech Recognition

ICASSP 2023accepted

Attention-based contextual biasing approaches have shown significant improvements in the recognition of generic and/or personal rare-words in End-to-End Automatic Speech Recognition (E2E ASR) systems like neural transducers. These approaches employ crossattention to bias the model towards specific c…

Cited by 0SourceScholar
2022

A Neural Prosody Encoder for End-to-End Dialogue Act Classification

ICASSP 2022accepted

Dialogue act classification (DAC) is a critical task for spoken language understanding in dialogue systems. Prosodic features such as energy and pitch have been shown to be useful for DAC. Despite their importance, little research has explored neural approaches to integrate prosodic features into en…

Cited by 0SourceScholar
2022

Caching Networks: Capitalizing on Common Speech for ASR

ICASSP 2022accepted

We introduce Caching Networks (CachingNets), a speech recognition network architecture capable of delivering faster, more accurate decoding by leveraging common speech patterns. By explicitly incorporating select sentences unique to each user into the network’s design, we show how to train the model…

Cited by 0SourceScholar
2022

Contextual Adapters for Personalized Speech Recognition in Neural Transducers

ICASSP 2022accepted

Personal rare word recognition in end-to-end Automatic Speech Recognition (E2E ASR) models is a challenge due to the lack of training data. A standard way to address this issue is with shallow fusion methods at inference time. However, due to their dependence on external language models and the dete…

Cited by 0SourceScholar
2022

Multi-Task RNN-T with Semantic Decoder for Streamable Spoken Language Understanding

ICASSP 2022accepted

End-to-end Spoken Language Understanding (E2E SLU) has attracted increasing interest due to its advantages of joint optimization and low latency when compared to traditionally cascaded pipelines. Existing E2E SLU models usually follow a two-stage configuration where an Automatic Speech Recognition (…

Cited by 0SourceScholar
2022

TINYS2I: A Small-Footprint Utterance Classification Model with Contextual Support for On-Device SLU

ICASSP 2022accepted

On-device spoken language understanding (SLU) offers the potential for significant latency savings compared to cloud-based processing, as the audio stream does not need to be transmitted to a server. We present Tiny Signal-to-interpretation (TinyS2I), an end-to-end on-device SLU approach which is fo…

Cited by 0SourceScholar
2021

Bifocal Neural ASR: Exploiting Keyword Spotting for Inference Optimization

ICASSP 2021accepted

We present Bifocal RNN-T, a new variant of the Recurrent Neural Network Transducer (RNN-T) architecture designed for improved inference time latency on speech recognition tasks. The architecture enables a dynamic pivot for its runtime compute pathway, namely taking advantage of keyword spotting to s…

Cited by 0SourceScholar