← Search

Kartik Audhkhasi

22 accepted papers

2025

Audio Diffusion with Large Language Models

ICASSP 2025accepted

In this paper, we explore an alternate approach to the popular method of using large language models (LLMs) as a second decoder for Automated Speech Recognition (ASR) and speech understanding tasks. We propose to employ diffusion networks to generate a correction signal that can be applied on the or…

Cited by 0SourceScholar
2025

Identifying and Mitigating Mismatched Language Code in Multilingual ASR

ICASSP 2025accepted

Multilingual speech recognition systems often use an input language code in order to prompt the transcription in the target language. However, the spoken language in the input audio may not always match the language code, as often prevalent in multilingual societies. This language mismatch can signi…

Cited by 0SourceScholar
2025

LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors

EMNLP 2025

Recently, large-scale pre-trained speech encoders and Large Language Models (LLMs) have been released, which show state-of-the-art performance on a range of spoken language processing tasks, including Automatic Speech Recognition (ASR). To effectively combine both models for better performance, cont

Cited by 0SourcePDFScholar
2025

Weak-to-Strong Generalization in Speech Recognition

ICASSP 2025accepted

To surpass human-level accuracy, speech recognition models must go beyond relying solely on human labels. To this end, we must build stronger models from weaker supervisors and this is the main goal in weak-to-strong generalization (WSG). WSG methods normally incorporate additional information into…

Cited by 0SourceScholar
2023

Large-Scale Language Model Rescoring on Long-Form Data

ICASSP 2023accepted

In this work, we study the impact of Large-scale Language Models (LLM) on Automated Speech Recognition (ASR) of YouTube videos, which we use as a source for long-form ASR. We demonstrate up to 8% relative reduction in Word Error Eate (WER) on US English (en-us) and code-switched Indian English (en-i…

Cited by 27SourceScholar
2023

Modular Conformer Training for Flexible End-to-End ASR

ICASSP 2023accepted

The state-of-the-art conformer used in automatic speech recognition combines feed-forward, convolution and multi-headed self-attention layers in a single model that is trained end-to-end with a decoder network. While this end-to-end training is simple and beneficial for word error rate, it restricts…

Cited by 0SourceScholar
2023

Robust Knowledge Distillation from RNN-T Models with Noisy Training Labels Using Full-Sum Loss

ICASSP 2023accepted

This work studies knowledge distillation (KD) and addresses its constraints for recurrent neural network transducer (RNN-T) models. In hard distillation, a teacher model transcribes large amounts of unlabelled speech to train a student model. Soft distillation is another popular KD method that disti…

Cited by 0SourceScholar
2021

Convolutional Dropout and Wordpiece Augmentation for End-to-End Speech Recognition

ICASSP 2021accepted

Regularization and data augmentation are crucial to training end-to-end automatic speech recognition systems. Dropout is a popular regularization technique, which operates on each neuron independently by multiplying it with a Bernoulli random variable. We propose a generalization of dropout, called…

Cited by 0SourceScholar
2020

Leveraging Unpaired Text Data for Training End-To-End Speech-to-Intent Systems

ICASSP 2020accepted

Training an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is time consuming and expensive to collect. Initializing the S2I model with an ASR model trained on copious speech data can…

Cited by 0SourceScholar
2019

Acoustically Grounded Word Embeddings for Improved Acoustics-to-word Speech Recognition

ICASSP 2019accepted

Direct acoustics-to-word (A2W) systems for end-to-end automatic speech recognition are simpler to train, and more efficient to decode with, than sub-word systems. However, A2W systems can have difficulties at training time when data is limited, and at decoding time when recognizing words outside the…

Cited by 0SourceScholar
2019

Sequence Noise Injected Training for End-to-end Speech Recognition

ICASSP 2019accepted

We present a simple noise injection algorithm for training end-to-end ASR models which consists in adding to the spectra of training utterances the scaled spectra of random utterances of comparable length. We conjecture that the sequence information of the "noise" utterances is important and verify…

Cited by 0SourceScholar
2018

Building Competitive Direct Acoustics-to-Word Models for English Conversational Speech Recognition

ICASSP 2018accepted

Direct acoustics-to-word (A2W) models in the end-to-end paradigm have received increasing attention compared to conventional subword based automatic speech recognition models using phones, characters, or context-dependent hidden Markov model states. This is because A2W models recognize words from sp…

Cited by 0SourceScholar
2018

Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition

ICASSP 2018accepted

The performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significant source of such mismatch. The traditional approach to deal with multiple accents involves pooling data from several ac…

Cited by 0SourceScholar
2017

End-to-end ASR-free keyword search from speech

ICASSP 2017accepted

End-to-end (E2E) systems have achieved competitive results compared to conventional hybrid hidden Markov model (HMM)-deep neural network based automatic speech recognition (ASR) systems. Such E2E systems are attractive due to the lack of dependence on alignments between input acoustic and output gra…

Cited by 0SourceScholar
2017

End-to-end speech recognition and keyword search on low-resource languages

ICASSP 2017accepted

In recent years, so-called, “end-to-end” speech recognition systems have emerged as viable alternatives to traditional ASR frameworks. Keyword search, localizing an orthographic query in a speech corpus, is typically performed by using automatic speech recognition (ASR) to generate an index. Previou…

Cited by 58SourceScholar
2017

Knowledge distillation across ensembles of multilingual models for low-resource languages

ICASSP 2017accepted

This paper investigates the effectiveness of knowledge distillation in the context of multilingual models. We show that with knowledge distillation, Long Short-Term Memory(LSTM) models can be used to train standard feed-forward Deep Neural Network (DNN) models for a variety of low-resource languages…

Cited by 0SourceScholar
2016

Efficient one-vs-one kernel ridge regression for speech recognition

ICASSP 2016accepted

Recent evidences suggest that the performance of kernel methods may match that of deep neural networks (DNNs), which have been the state-of-the-art approach for speech recognition. In this work, we present an improvement of the kernel ridge regression studied in Huang et al., ICASSP 2014, and show t…

Cited by 0SourceScholar
2016

Semantic word embedding neural network language models for automatic speech recognition

ICASSP 2016accepted

Semantic word embeddings have become increasingly important in natural language processing tasks over the last few years. This popularity is due to their ability to easily capture rich semantic information through a distributed representation and the availability of fast and scalable algorithms for…

Cited by 0SourceScholar
2015

A mixture of experts approach towards intelligibility classification of pathological speech

ICASSP 2015accepted

Pathological speech involves atypical speech production which may result from several factors including oral diseases, physical disabilities in the voice production system and atypical anatomy. Automatic evaluation of intelligibility in patients with pathological speech can assist accurate diagnosis…

Cited by 0SourceScholar