← Search

Aravind Ganapathiraju

7 accepted papers

2025

Speech Data Selection for Efficient ASR Fine-Tuning using Domain Classifier and Pseudo-Label Filtering

ICASSP 2025accepted

In real-world speech data processing, the scarcity of annotated data and the abundance of unlabelled speech data present a significant challenge. To address this, we propose an efficient data selection pipeline for fine-tuning ASR models by generating pseudo-labels using WhisperX pipeline and select…

Cited by 0SourceScholar
2025

XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models

ICASSP 2025accepted

Self-supervised pretrained models exhibit competitive performance in automatic speech recognition (ASR) on finetuning, even with limited in-domain supervised data. However, popular pretrained models are not suitable for streaming ASR because they are trained with full attention context. In this pape…

Cited by 0SourceScholar
2024

Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper

EMNLP 2024finding

The training of automatic speech recognition (ASR) with little to no supervised data remains an open question. In this work, we demonstrate that streaming Transformer-Transducer (TT) models can be trained from scratch in consumer and accessible GPUs in their entirety with pseudo-labeled (PL) speech…

2024

Multitask Speech Recognition and Speaker Change Detection for Unknown Number of Speakers

ICASSP 2024accepted

Traditionally, automatic speech recognition (ASR) and speaker change detection (SCD) systems have been independently trained to generate comprehensive transcripts accompanied by speaker turns. Recently, joint training of ASR and SCD systems, by inserting speaker turn tokens in the ASR training text,…

Cited by 0SourceScholar
2024

Probability-Aware Word-Confusion-Network-To-Text Alignment Approach for Intent Classification

ICASSP 2024accepted

Spoken Language Understanding (SLU) technologies have greatly improved due to the effective pretraining of speech representations. A common requirement of industry-based solutions is the portability to deploy SLU models in voice-assistant devices. Thus, distilling knowledge from large text-based lan…

Cited by 0SourceScholar
2024

TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR

EMNLP 2024main

In traditional conversational intelligence from speech, a cascaded pipeline is used, involving tasks such as voice activity detection, diarization, transcription, and subsequent processing with different NLP models for tasks like semantic endpointing and named entity recognition (NER). Our paper int…

2023

Effectiveness of Text, Acoustic, and Lattice-Based Representations in Spoken Language Understanding Tasks

ICASSP 2023accepted

In this paper, we perform an exhaustive evaluation of different representations to address the intent classification problem in a Spoken Language Understanding (SLU) setup. We benchmark three types of systems to perform the SLU intent detection task: 1) text-based, 2) lattice-based, and a novel 3) m…

Cited by 0SourceScholar