← Search

Anirudh Raju

8 accepted papers

2024

Hot-Fixing Wake Word Recognition for End-to-End ASR Via Neural Model Reprogramming

ICASSP 2024accepted

This paper proposes two novel variants of neural reprogramming to enhance wake word recognition in streaming end-to-end ASR models without updating model weights. The first, "trigger-frame reprogramming", prepends the input speech feature sequence with the learned trigger-frames of the target wake w…

Cited by 4SourceScholar
2024

Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion

ICASSP 2024accepted

We propose an approach for continuous prediction of turn-taking and backchanneling locations in spoken dialogue by fusing a neural acoustic model with a large language model (LLM). Experiments on the Switchboard human-human conversation dataset demonstrate that our approach consistently outperforms…

Cited by 26SourceScholar
2023

Adaptive Endpointing with Deep Contextual Multi-Armed Bandits

ICASSP 2023accepted

Current endpointing (EP) solutions learn in a supervised framework, which does not allow the model to incorporate feedback and improve in an online setting. Also, it is common practice to utilize costly grid-search to find the best configuration for an endpointing model. In this paper, we aim to pro…

Cited by 0SourceScholar
2023

Cross-Utterance ASR Rescoring with Graph-Based Label Propagation

ICASSP 2023accepted

We propose a novel approach for ASR N-best hypothesis rescoring with graph-based label propagation by leveraging cross-utterance acoustic similarity. In contrast to conventional neural language model (LM) based ASR rescoring/reranking models, our approach focuses on acoustic information and conducts…

Cited by 0SourceScholar
2023

Federated Self-Learning with Weak Supervision for Speech Recognition

ICASSP 2023accepted

Automatic speech recognition (ASR) models with low-footprint are increasingly being deployed on edge devices for conversational agents, which enhances privacy. We study the problem of federated continual incremental learning for recurrent neural network-transducer (RNN-T) ASR models in the privacy-e…

Cited by 0SourceScholar
2021

DO as I Mean, Not as I Say: Sequence Loss Training for Spoken Language Understanding

ICASSP 2021accepted

Spoken language understanding (SLU) systems extract transcriptions, as well as semantics of intent or named entities from speech, and are essential components of voice activated systems. SLU models, which either directly extract semantics from audio or are composed of pipelined automatic speech reco…

Cited by 0SourceScholar
2019

Improving Noise Robustness of Automatic Speech Recognition via Parallel Data and Teacher-student Learning

ICASSP 2019accepted

For real-world speech recognition applications, noise robustness is still a challenge. In this work, we adopt the teacher-student (T/S) learning technique using a parallel clean and noisy corpus for improving automatic speech recognition (ASR) performance under multimedia noise. On top of that, we a…

Cited by 53SourceScholar
2018

Time-Delayed Bottleneck Highway Networks Using a DFT Feature for Keyword Spotting

ICASSP 2018accepted

This paper presents a novel deep neural network (DNN) architecture with highway blocks (HWs) using a complex discrete Fourier transform (DFT) feature for keyword spotting. In our previous work, we showed that the feed-forward DNN with a time-delayed bottleneck layer (TDB-DNN) directly trained from t…

Cited by 0SourceScholar