← Search

Sanjeev Khudanpur

45 accepted papers

2026

SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper

ICASSP 2026oral

Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a major challenge. While some approaches achieve strong performance when fine-tuned on specific domains, few systems generalize well across out-of-domain datasets. Our prior work, Diarization-Conditioned Whis…

Cited by 4SourcePDFScholar
2025

Benchmarking Language Model Creativity: A Case Study on Code Generation

NAACL 2025long

As LLMs become increasingly prevalent, it is interesting to consider how “creative” these models can be. From cognitive science, creativity consists of at least two key characteristics: convergent thinking (purposefulness to achieve a given goal) and divergent thinking (adaptability to explore new e…

2025

HLTCOE Submission to the VoicePrivacy Attacker Challenge

ICASSP 2025accepted

We describe our submission to the 2024 VoicePrivacy Attacker Challenge. We propose three main categories of methods to improve ASV performance against anonymized speech: improvements to the underlying classifier, alternative distance metrics when computing ASV scores, and kNN-VC normalization. By si…

Cited by 0SourceScholar
2025

Target Speaker ASR with Whisper

ICASSP 2025accepted

We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier to model relative differences among speakers by learning to condition on frame-level diarization outputs than to learn th…

Cited by 0SourceScholar
2025

Whisper-UT: A Unified Translation Framework for Speech and Text

EMNLP 2025

Encoder-decoder models have achieved remarkable success in speech and text tasks, yet efficiently adapting these models to diverse uni/multi-modal scenarios remains an open challenge. In this paper, we propose Whisper-UT, a unified and efficient framework that leverages lightweight adapters to enabl

Cited by 0SourcePDFScholar
2024

ConEC: Earnings Call Dataset with Real-world Contexts for Benchmarking Contextual Speech Recognition

COLING 2024main

Knowing the particular context associated with a conversation can help improving the performance of an automatic speech recognition (ASR) system. For example, if we are provided with a list of in-context words or phrases — such as the speaker’s contacts or recent song playlists — during inference, w…

2024

Enhancing Code-Switching Speech Recognition With Interactive Language Biases

ICASSP 2024accepted

Languages usually switch within a multilingual speech signal, especially in a bilingual society. This phenomenon is referred to as code-switching (CS), making automatic speech recognition (ASR) challenging under a multilingual scenario. We propose to improve CS-ASR by biasing the hybrid CTC/attentio…

Cited by 30SourceScholar
2024

Enhancing End-to-End Conversational Speech Translation Through Target Language Context Utilization

ICASSP 2024accepted

Incorporating longer context has been shown to benefit machine translation, but the inclusion of context in end-to-end speech translation (E2E-ST) remains under-studied. To bridge this gap, we introduce target language context in E2E-ST, enhancing coherence and overcoming memory constraints of exten…

Cited by 0SourceScholar
2024

Kreyòl-MT: Building MT for Latin American, Caribbean and Colonial African Creole Languages

NAACL 2024long

A majority of language technologies are tailored for a small number of high-resource languages, while relatively many low-resource languages are neglected. One such group, Creole languages, have long been marginalized in academic study, though their speakers could benefit from machine translation (M…

2024

Less Peaky and More Accurate CTC Forced Alignment by Label Priors

ICASSP 2024accepted

Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can cause inaccurate forced alignments (FA), especially at finer granularity, e.g., phoneme level. This paper aims at allevia…

Cited by 0SourceScholar
2024

Speech Collage: Code-Switched Audio Generation by Collaging Monolingual Corpora

ICASSP 2024accepted

Designing effective automatic speech recognition (ASR) systems for Code-Switching (CS) often depends on the availability of the transcribed CS resources. To address data scarcity, this paper introduces Speech Collage, a method that synthesizes CS data from monolingual corpora by splicing audio segme…

Cited by 0SourceScholar
2023

Adapting Self-Supervised Models to Multi-Talker Speech Recognition Using Speaker Embeddings

ICASSP 2023accepted

Self-supervised learning (SSL) methods which learn representations of data without explicit supervision have gained popularity in speech-processing tasks, particularly for single-talker applications. However, these models often have degraded performance for multi-talker scenarios — possibly due to t…

Cited by 45SourceScholar
2023

Building Keyword Search System from End-To-End Asr Systems

ICASSP 2023accepted

Keyword search (KWS) systems are commonly built on top of existing automatic speech recognition (ASR) systems. However, end-to-end (E2E) ASR models are not naturally equipped with word-level timing information or confidence. Existing methods for re-purposing E2E ASR systems for KWS are largely heuri…

Cited by 0SourceScholar
2023

Euro: Espnet Unsupervised ASR Open-Source Toolkit

ICASSP 2023accepted

This paper describes the ESPnet Unsupervised ASR Open-source Toolkit (EURO), an end-to-end open-source toolkit for unsupervised automatic speech recognition (UASR). EURO adopts the state-of-the-art UASR learning method introduced by the Wav2vec-U, originally implemented at FAIRSEQ, which leverages s…

Cited by 0SourceScholar
2023

Reducing Language Confusion for Code-Switching Speech Recognition with Token-Level Language Diarization

ICASSP 2023accepted

Code-switching (CS) occurs when languages switch within a speech signal and leads to language confusion for automatic speech recognition (ASR). We address the problem of language confusion for improving CS-ASR from two perspectives: incorporating and disentangling language information. We incorporat…

Cited by 0SourceScholar
2022

Injecting Text and Cross-Lingual Supervision in Few-Shot Learning from Self-Supervised Models

ICASSP 2022accepted

Self-supervised model pretraining has recently garnered significant interest. However, using additional resources in fine-tuning these models has received less attention. We demonstrate how universal phoneset acoustic models can leverage cross-lingual supervision to improve transfer of pretrained se…

Cited by 0SourceScholar
2022

Investigating Self-Supervised Learning for Speech Enhancement and Separation

ICASSP 2022accepted

Speech enhancement and separation are two fundamental tasks for robust speech processing. Speech enhancement suppresses background noise while speech separation extracts target speech from interfering speakers. Despite a great number of supervised learning-based enhancement and separation methods ha…

Cited by 0SourceScholar
2021

An Asynchronous WFST-Based Decoder for Automatic Speech Recognition

ICASSP 2021accepted

We introduce asynchronous dynamic decoder, which adopts an efficient A* algorithm to incorporate big language models in the one-pass decoding for large vocabulary continuous speech recognition. Unlike standard one-pass decoding with on-the-fly composition decoder which might induce a significant com…

Cited by 0SourceScholar
2021

Fine-Grained Activity Recognition for Assembly Videos

RA-L 2021

In this letter we address the task of recognizing assembly actions as a structure (e.g. a piece of furniture or a toy block tower) is built up from a set of primitive objects. Recognizing the full range of assembly actions requires perception at a level of spatial detail that has not been attempted

Cited by 19SourceScholar
2021

Training Noisy Single-Channel Speech Separation with Noisy Oracle Sources: A Large Gap and a Small Step

ICASSP 2021accepted

As the performance of single-channel speech separation systems has improved, there has been a desire to move to more challenging conditions than the clean, near-field speech that initial systems were developed on. When training deep learning separation models, a need for ground truth leads to traini…

Cited by 0SourceScholar
2020

An Empirical Study of Transformer-Based Neural Language Model Adaptation

ICASSP 2020accepted

We explore two adaptation approaches of deep Transformer based neural language models (LMs) for automatic speech recognition. The first approach is a pretrain-finetune framework, where we first pretrain a Transformer LM on a large-scale text corpus from scratch and then adapt it to relatively small…

Cited by 32SourceScholar
2020

OOV Recovery with Efficient 2nd Pass Decoding and Open-vocabulary Word-level RNNLM Rescoring for Hybrid ASR

ICASSP 2020accepted

In this paper, we investigate out-of-vocabulary (OOV) word recovery in hybrid automatic speech recognition (ASR) systems, with emphasis on dynamic vocabulary expansion for both Weight Finite State Transducer (WFST)-based decoding and word-level RNNLM rescoring. We first describe our OOV candidate ge…

Cited by 0SourceScholar
2020

Speaker Diarization with Region Proposal Network

ICASSP 2020accepted

Speaker diarization is an important pre-processing step for many speech applications, and it aims to solve the "who spoke when" problem. Although the standard diarization systems can achieve satisfactory results in various scenarios, they are composed of several independently-optimized modules and c…

Cited by 62SourceScholar
2019

Acoustic Modeling for Overlapping Speech Recognition: Jhu Chime-5 Challenge System

ICASSP 2019accepted

This paper summarizes our acoustic modeling efforts in the Johns Hopkins University speech recognition system for the CHiME-5 challenge to recognize highly-overlapped dinner party speech recorded by multiple microphone arrays. We explore data augmentation approaches, neural network architectures, fr…

Cited by 0SourceScholar
2019

Speaker Recognition for Multi-speaker Conversations Using X-vectors

ICASSP 2019accepted

Recently, deep neural networks that map utterances to fixed-dimensional embeddings have emerged as the state-of-the-art in speaker recognition. Our prior work introduced x-vectors, an embedding that is very effective for both speaker recognition and diarization. This paper combines our previous work…

Cited by 0SourceScholar
2018

A Pruned Rnnlm Lattice-Rescoring Algorithm for Automatic Speech Recognition

ICASSP 2018accepted

Lattice-rescoring is a common approach to take advantage of recurrent neural language models in ASR, where a word-lattice is generated from 1st-pass decoding and the lattice is then rescored with a neural model, and an <i xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/…

Cited by 0SourceScholar
2018

A Time-Restricted Self-Attention Layer for ASR

ICASSP 2018accepted

Self-attention - an attention mechanism where the input and output sequence lengths are the same - has recently been successfully applied to machine translation, caption generation, and phoneme recognition. In this paper we apply a restricted self-attention mechanism (with multiple heads) to speech…

Cited by 0SourceScholar
2018

Bayesian Models for Unit Discovery on a Very Low Resource Language

ICASSP 2018accepted

Developing speech technologies for low-resource languages has become a very active research field over the last decade. Among others, Bayesian models have shown some promising results on artificial examples but still lack of in situ experiments. Our work applies state-of-the-art Bayesian models to u…

Cited by 0SourceScholar
2018

Characterizing Performance of Speaker Diarization Systems on Far-Field Speech Using Standard Methods

ICASSP 2018accepted

To date, the bulk of research on speaker diarization has been conducted on telephone or near-field speech. As the need for technologies capable of handling conversational speech increases, it is necessary to establish the performance of state-of-the-art systems in this domain. In this work we evalua…

Cited by 0SourceScholar
2018

Enhancement and Analysis of Conversational Speech: JSALT 2017

ICASSP 2018accepted

Automatic speech recognition is more and more widely and effectively used. Nevertheless, in some automatic speech analysis tasks the state of the art is surprisingly poor. One of these is “diarization”, the task of determining who spoke when. Diarization is key to processing meeting audio and clinic…

Cited by 0SourceScholar
2018

Neural Network Language Modeling with Letter-Based Features and Importance Sampling

ICASSP 2018accepted

In this paper we describe an extension of the Kaldi software toolkit to support neural-based language modeling, intended for use in automatic speech recognition (ASR) and related tasks. We combine the use of subword features (letter n-grams) and one-hot encoding of frequent words so that the models…

Cited by 0SourceScholar
2018

Semi-Supervised Training of Acoustic Models Using Lattice-Free MMI

ICASSP 2018accepted

The lattice-free MMI objective (LF-MMI) has been used in supervised training of state-of-the-art neural network acoustic models for automatic speech recognition (ASR). With large amounts of unsupervised data available, extending this approach to the semi-supervised scenario is of significance. Finit…

Cited by 0SourceScholar
2018

X-Vectors: Robust DNN Embeddings for Speaker Recognition

ICASSP 2018accepted

In this paper, we use data augmentation to improve performance of deep neural network (DNN) embeddings for speaker recognition. The DNN, which is trained to discriminate between speakers, maps variable-length utterances to fixed-dimensional embeddings that we call x-vectors. Prior studies have found…

Cited by 0SourceScholar
2017

A study on data augmentation of reverberant speech for robust speech recognition

ICASSP 2017accepted

The environmental robustness of DNN-based acoustic models can be significantly improved by using multi-condition training data. However, as data collection is a costly proposition, simulation of the desired conditions is a frequently adopted strategy. In this paper we detail a data augmentation appr…

Cited by 0SourceScholar
2017

An empirical evaluation of zero resource acoustic unit discovery

ICASSP 2017accepted

Acoustic unit discovery (AUD) is a process of automatically identifying a categorical acoustic unit inventory from speech and producing corresponding acoustic unit tokenizations. AUD provides an important avenue for unsupervised acoustic model training in a zero resource setting where expert-provide…

Cited by 0SourceScholar
2017

Topic identification of spoken documents using unsupervised acoustic unit discovery

ICASSP 2017accepted

This paper investigates the application of unsupervised acoustic unit discovery for topic identification (topic ID) of spoken audio documents. The acoustic unit discovery method is based on a non-parametric Bayesian phone-loop model that segments a speech utterance into phone-like categories. The di…

Cited by 0SourceScholar
2016

Acoustic data-driven pronunciation lexicon generation for logographic languages

ICASSP 2016accepted

Handcrafted pronunciation lexicons are widely used in modern speech recognition systems. Designing a pronunciation lexicon, however, requires tremendous amount of expert knowledge and effort, which is not practical when applying speech recognition techniques to low resource languages. In this paper,…

Cited by 0SourceScholar
2016

Adapting ASR for under-resourced languages using mismatched transcriptions

ICASSP 2016accepted

Mismatched transcriptions of speech in a target language refers to transcriptions provided by people unfamiliar with the language, using English letter sequences. In this work, we demonstrate the value of such transcriptions in building an ASR system for the target language. For different languages,…

Cited by 0SourceScholar
2016

Context-dependent point process models for keyword search and detection-based ASR

ICASSP 2016accepted

The point process model (PPM) for keyword search (KWS) is a whole-word parametric approach that characterizes each query type by the timing of phonetic events observed during its production. In this paper, we first extend the PPM modeling framework to operate on context-dependent phonetic event patt…

Cited by 0SourceScholar
2016

Highway long short-term memory RNNS for distant speech recognition

ICASSP 2016accepted

In this paper, we extend the deep long short-term memory (DL-STM) recurrent neural networks by introducing gated direct connections between memory cells in adjacent layers. These direct links, called highway connections, enable unimpeded information flow across different layers and thus alleviate th…

Cited by 0SourceScholar
2016

Unsupervised surgical data alignment with application to automatic activity annotation

ICRA 2016

Robotic surgery and other minimally-invasive surgical techniques are an integral part of patient care, and readily yield large amounts of data. Surgical tool motion (kinematic data) contains information that is useful for assessment and education. Typically, assessment and education tools that rely

Cited by 16SourceScholar
2015

Librispeech: An ASR corpus based on public domain audio books

ICASSP 2015accepted

This paper introduces a new corpus of read English speech, suitable for training and evaluating speech recognition systems. The LibriSpeech corpus is derived from audiobooks that are part of the LibriVox project, and contains 1000 hours of speech sampled at 16 kHz. We have made the corpus freely ava…

Cited by 0SourceScholar
2015

Towards machines that know when they do not know: Summary of work done at 2014 Frederick Jelinek Memorial Workshop

ICASSP 2015accepted

A group of junior and senior researchers gathered as a part of the 2014 Frederick Jelinek Memorial Workshop in Prague to address the problem of predicting the accuracy of a nonlinear Deep Neural Network probability estimator for unknown data in a different application domain from the domain in which…

Cited by 0SourceScholar