← Search

Kaustubh Kalgaonkar

8 accepted papers

2025

Directional Source Separation for Robust Speech Recognition on Smart Glasses

ICASSP 2025accepted

Modern smart glasses leverage machine learning to offer real-time transcriptions, considerably enriching human communication experiences. However, such systems frequently encounter challenges related to environmental noises, leading to decreased speech recognition. To improve voice quality, this wor…

Cited by 15SourceScholar
2023

Egocentric Audio-Visual Noise Suppression

ICASSP 2023accepted

This paper studies audio-visual noise suppression for egocentric videos -where the speaker is not captured in the video. Instead, potential noise sources are visible on screen with the camera emulating the off-screen speaker’s view of the outside world. This setting is different from prior work in a…

Cited by 0SourceScholar
2023

SCA: Streaming Cross-Attention Alignment For Echo Cancellation

ICASSP 2023accepted

End-to-End deep learning has shown promising results for speech enhancement tasks, such as noise suppression, dereverberation, and speech separation. However, most state-of-the-art methods for echo cancellation are either classical DSP-based or hybrid DSP-ML algorithms. Components such as the delay…

Cited by 0SourceScholar
2022

Architecture for Variable Bitrate Neural Speech Codec with Configurable Computation Complexity

ICASSP 2022accepted

Low bitrate speech codecs have become an area of intense research. Traditional speech codecs, which use signal processing methods to encode and decode speech, often suffer from quality issues at low bitrates. A neural speech codec, which uses a deep neural network in the compression pipeline, can he…

Cited by 0SourceScholar
2021

A Time-Domain Convolutional Recurrent Network for Packet Loss Concealment

ICASSP 2021accepted

Packet loss may affect a wide range of applications that use voice over IP (VoIP), e.g. video conferencing. In this paper, we investigate a time-domain convolutional recurrent network (CRN) for online packet loss concealment. The CRN comprises a convolutional encoder-decoder structure and long short…

Cited by 0SourceScholar
2020

Spatial Attention for Far-Field Speech Recognition with Deep Beamforming Neural Networks

ICASSP 2020accepted

In this paper, we introduce spatial attention for refining the information in multi-direction neural beamformer for far-field automatic speech recognition. Previous approaches of neural beamformers with multiple look directions, such as the factored complex linear projection, have shown promising re…

Cited by 0SourceScholar
2015

Estimating confidence scores on ASR results using recurrent neural networks

ICASSP 2015accepted

In this paper we present a confidence estimation system using recurrent neural networks (RNN) and compare it to a traditional multilayered perception (MLP) based system. The ability of RNN to capture sequence information and improve decisions using processed history was main motivation to explore RN…

Cited by 0SourceScholar