← Search

Sameer Khurana

14 accepted papers

2025

Aligning Multimodal Representations through an Information Bottleneck

ICML 2025poster

Contrastive losses have been extensively used as a tool for multimodal representation learning. However, it has been empirically observed that their use is not effective to learn an aligned representation space. In this paper, we argue that this phenomenon is caused by the presence of modality-spec…

Cited by 0SourcePDFScholar
2025

Interactive Robot Action Replanning using Multimodal LLM Trained from Human Demonstration Videos

ICASSP 2025accepted

Understanding human actions could allow robots to perform a large spectrum of complex manipulation tasks and make collaboration with humans easier. Recently, multimodal scene understanding using audio-visual Transformers has been used to generate robot action sequences from videos of human demonstra…

Cited by 0SourceScholar
2025

Leveraging Audio-Only Data for Text-Queried Target Sound Extraction

ICASSP 2025accepted

The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access to large-scale text-audio pairs to address a variety of text queries, the limited number of available high-quality text-…

Cited by 0SourceScholar
2024

Generation or Replication: Auscultating Audio Latent Diffusion Models

ICASSP 2024accepted

The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt at understanding the inner workings of audio latent diffusi…

Cited by 0SourceScholar
2024

NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization

ICASSP 2024accepted

Head-related transfer functions (HRTFs) are important for immersive audio, and their spatial interpolation has been studied to upsample finite measurements. Recently, neural fields (NFs) which map from sound source direction to HRTF have gained attention. Existing NF-based methods focused on estimat…

Cited by 0SourceScholar
2024

NeuroHeed+: Improving Neuro-Steered Speaker Extraction with Joint Auditory Attention Detection

ICASSP 2024accepted

Neuro-steered speaker extraction aims to extract the listener’s brainattended speech signal from a multi-talker speech signal, in which the attention is derived from the cortical activity. This activity is usually recorded using electroencephalography (EEG) devices. Though promising, current methods…

Cited by 0SourceScholar
2024

WI-FI based Indoor Monitoring Enhanced by Multimodal Fusion

ICASSP 2024accepted

Indoor monitoring systems are in high demand to protect vulnerable people, especially when they are alone at home, in nursing homes, hospitals, etc. Although surveillance systems in public spaces use cameras and microphones to find incidents, indoor monitoring in personal spaces needs to protect pri…

Cited by 0SourceScholar
2023

On Unsupervised Uncertainty-Driven Speech Pseudo-Label Filtering and Model Calibration

ICASSP 2023accepted

Pseudo-label (PL) filtering forms a crucial part of Self-Training (ST) methods for unsupervised domain adaptation. Dropout-based Uncertainty-driven Self-Training (DUST) proceeds by first training a teacher model on source domain labeled data. Then, the teacher model is used to provide PLs for the un…

Cited by 0SourceScholar
2022

Detecting Dementia from Long Neuropsychological Interviews

EMNLP 2022finding

Neuropsychological exams are commonly used to diagnose various kinds of cognitive impairment. They typically involve a trained examiner who conducts a series of cognitive tests with a subject. In recent years, there has been growing interest in developing machine learning methods to extract speech a…

Cited by 3SourcePDFScholar
2021

PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition

NeurIPS 2021spotlight

Self-supervised speech representation learning (speech SSL) has demonstrated the benefit of scale in learning rich representations for Automatic Speech Recognition (ASR) with limited paired data, such as wav2vec 2.0. We investigate the existence of sparse subnetworks in pre-trained speech SSL models…

Cited by 80SourcePDFScholar
2021

Unsupervised Domain Adaptation for Speech Recognition via Uncertainty Driven Self-Training

ICASSP 2021accepted

The performance of automatic speech recognition (ASR) systems typically degrades significantly when the training and test data domains are mismatched. In this paper, we show that self-training (ST) combined with an uncertainty-based pseudo-label filtering approach can be effectively used for domain…

Cited by 0SourceScholar
2019

A Factorial Deep Markov Model for Unsupervised Disentangled Representation Learning from Speech

ICASSP 2019accepted

We present the Factorial Deep Markov Model (FDMM) for representation learning of speech. The FDMM learns disentangled, interpretable and lower dimensional latent representations from speech without supervision. We use a static and dynamic latent variable to exploit the fact that information in a spe…

Cited by 0SourceScholar
2018

Exploiting Convolutional Neural Networks for Phonotactic Based Dialect Identification

ICASSP 2018accepted

In this paper, we investigate different approaches for Dialect Identification (DID) in Arabic broadcast speech. Dialects differ in their inventory of phonological segments. This paper proposes a new phonotactic based feature representation approach which enables discrimination among different occurr…

Cited by 0SourceScholar