← Search

Pranav Dheram

4 accepted papers

2024

Hot-Fixing Wake Word Recognition for End-to-End ASR Via Neural Model Reprogramming

ICASSP 2024accepted

This paper proposes two novel variants of neural reprogramming to enhance wake word recognition in streaming end-to-end ASR models without updating model weights. The first, "trigger-frame reprogramming", prepends the input speech feature sequence with the learned trigger-frames of the target wake w…

Cited by 4SourceScholar
2024

Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition

COLING 2024main

Recent advances in machine learning have demonstrated that multi-modal pre-training can improve automatic speech recognition (ASR) performance compared to randomly initialized models, even when models are fine-tuned on uni-modal tasks. Existing multi-modal pre-training methods for the ASR task have…

Cited by 2SourcePDFScholar
2024

Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion

ICASSP 2024accepted

We propose an approach for continuous prediction of turn-taking and backchanneling locations in spoken dialogue by fusing a neural acoustic model with a large language model (LLM). Experiments on the Switchboard human-human conversation dataset demonstrate that our approach consistently outperforms…

Cited by 26SourceScholar
2021

DO as I Mean, Not as I Say: Sequence Loss Training for Spoken Language Understanding

ICASSP 2021accepted

Spoken language understanding (SLU) systems extract transcriptions, as well as semantics of intent or named entities from speech, and are essential components of voice activated systems. SLU models, which either directly extract semantics from audio or are composed of pipelined automatic speech reco…

Cited by 0SourceScholar