← Search

Kshitiz Kumar

7 accepted papers

2022

Maximizing Audio Event Detection Model Performance on Small Datasets Through Knowledge Transfer, Data Augmentation, and Pretraining: an Ablation Study

ICASSP 2022accepted

An Xception model reaches state-of-the-art (SOTA) accuracy on the ESC-50 dataset for audio event detection through knowledge transfer from ImageNet weights, pretraining on AudioSet, and an on-the-fly data augmentation pipeline. This paper presents an ablation study that analyzes which components con…

Cited by 0SourceScholar
2019

Word Characters and Phone Pronunciation Embedding for ASR Confidence Classifier

ICASSP 2019accepted

Confidence classifier is an integral component of an automatic speech recognition (ASR) system. These classifiers predict the accuracy of an ASR hypothesis by associating a confidence score in [0,1] range, where larger score implies higher probability of the hypothesis being correct. Confidence scor…

Cited by 0SourceScholar
2017

Extended low-rank plus diagonal adaptation for deep and recurrent neural networks

ICASSP 2017accepted

Recently, the low-rank plus diagonal (LRPD) adaptation was proposed for speaker adaptation of deep neural network (DNN) models. The LRPD restructures the adaptation matrix as a superposition of a diagonal matrix and a product of two low-rank matrices. In this paper, we extend the LRPD adaptation int…

Cited by 0SourceScholar
2016

Investigations on speaker adaptation of LSTM RNN models for speech recognition

ICASSP 2016accepted

Recently Long Short-Term Memory (LSTM) Recurrent Neural Networks (RNN) acoustic models have demonstrated superior performance over deep neural networks (DNN) models in speech recognition and many other tasks. Although a lot of work have been reported on DNN model adaptation, very little has been don…

Cited by 0SourceScholar
2016

Non-negative intermediate-layer DNN adaptation for a 10-KB speaker adaptation profile

ICASSP 2016accepted

Previously we demonstrated that speaker adaptation of acoustic models (AM) can provide significant improvement in the accuracy of large-scale speech recognition systems. In this work we discuss numerous challenges in scaling speaker adaptation to millions of speakers, where the size of speaker-depen…

Cited by 0SourceScholar