← Search

Venkatesh Ravichandran

5 accepted papers

2024

Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition

COLING 2024main

Recent advances in machine learning have demonstrated that multi-modal pre-training can improve automatic speech recognition (ASR) performance compared to randomly initialized models, even when models are fine-tuned on uni-modal tasks. Existing multi-modal pre-training methods for the ASR task have…

Cited by 2SourcePDFScholar
2024

Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion

ICASSP 2024accepted

We propose an approach for continuous prediction of turn-taking and backchanneling locations in spoken dialogue by fusing a neural acoustic model with a large language model (LLM). Experiments on the Switchboard human-human conversation dataset demonstrate that our approach consistently outperforms…

Cited by 26SourceScholar
2023

Adaptive Endpointing with Deep Contextual Multi-Armed Bandits

ICASSP 2023accepted

Current endpointing (EP) solutions learn in a supervised framework, which does not allow the model to incorporate feedback and improve in an online setting. Also, it is common practice to utilize costly grid-search to find the best configuration for an endpointing model. In this paper, we aim to pro…

Cited by 0SourceScholar
2023

Cross-Utterance ASR Rescoring with Graph-Based Label Propagation

ICASSP 2023accepted

We propose a novel approach for ASR N-best hypothesis rescoring with graph-based label propagation by leveraging cross-utterance acoustic similarity. In contrast to conventional neural language model (LM) based ASR rescoring/reranking models, our approach focuses on acoustic information and conducts…

Cited by 0SourceScholar
2023

Towards Accurate and Real-Time End-of-Speech Estimation

ICASSP 2023accepted

We introduce a variant of the endpoint (EP) detection problem in automatic speech recognition (ASR), which we call the end-of-speech (EOS) estimation. Given an utterance, EOS estimation aims to identify the timestamp when the utterance waveform has fully decayed and is then used to measure the EP la…

Cited by 0SourceScholar