← Search

Anurag Kumar

40 accepted papers

2026

PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs

ICML 2026poster

Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained to fixed microphone geometries, preventing deployment across diverse devices. We present PhaseCoder, a transformer-only …

Cited by 0SourceScholar
2025

Advancing Active Speaker Detection for Egocentric Videos

ICASSP 2025accepted

This paper presents an improved approach to multimodal active speaker detection in egocentric videos, specifically designed to be robust against the rapid movements and motion blur commonly found in such videos. We propose two key techniques to improve the model’s resilience: (i) spatially fixing th…

Cited by 0SourceScholar
2025

Bridging Context Gaps: Enhancing Comprehension in Long-Form Social Conversations Through Contextualized Excerpts

COLING 2025main

We focus on enhancing comprehension in small-group recorded conversations, which serve as a medium to bring people together and provide a space for sharing personal stories and experiences on crucial social matters. One way to parse and convey information from these conversations is by sharing highl…

2025

Hearing Anywhere in Any Environment

CVPR 2025poster

In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR) estimation, most existing methods are limited to the single environ…

Cited by 0SourcePDFScholar
2025

Learning to Highlight Audio by Watching Movies

CVPR 2025poster

Recent years have seen a significant increase in video content creation and consumption. Crafting engaging content requires the careful curation of both visual and audio elements. While visual cue curation, through techniques like optimal viewpoint selection or post-editing, has been central to medi…

Cited by 0SourcePDFScholar
2025

Reexamining the Efficacy of MetricGAN for Speech Enhancement

ICASSP 2025accepted

MetricGAN, a notable generative approach, provides an effective framework to train speech enhancement models to produce high metric scores. However, we identify two key limitations of current MetricGAN-family models, i.e. neglecting certain mainstream metrics during evaluation and conducting evaluat…

Cited by 0SourceScholar
2025

SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models

ICASSP 2025accepted

Speaker Diarization (SD) is a crucial component of modern end-to-end ASR pipelines. Traditional SD systems, which are typically audio-based and operate independently of ASR, often introduce speaker errors, particularly during speaker transitions and overlapping speech. Recently, language models incl…

Cited by 0SourceScholar
2025

Using RLHF to align speech enhancement approaches to mean-opinion quality scores

ICASSP 2025accepted

Objective speech quality measures are typically used to assess speech enhancement algorithms, but it has been shown that they are sub-optimal as learning objectives because they do not always align well with human subjective ratings. This misalignment often results in noticeable distortions and arti…

Cited by 0SourceScholar
2024

A Closer Look at Wav2vec2 Embeddings for On-Device Single-Channel Speech Enhancement

ICASSP 2024accepted

Self-supervised learned models have been found to be very effective for tasks such as automatic speech recognition, speaker identification, and others. However, their utility in speech enhancement systems is yet to be firmly established, and perhaps slightly misunderstood. In this paper, we investig…

Cited by 0SourceScholar
2024

Ambisonics Networks - The Effect of Radial Functions Regularization

ICASSP 2024accepted

Ambisonics, a popular format of spatial audio, is the spherical harmonic (SH) representation of the plane wave density function of a sound field. Many algorithms operate in the SH domain and utilize the Ambisonics as their input signal. The process of encoding Ambisonics from a spherical microphone…

Cited by 0SourceScholar
2024

Audiovisual Speaker Separation with Full- and Sub-Band Modeling in the Time-Frequency Domain

ICASSP 2024accepted

We introduce a new deep learning model for talker-independent audiovisual speaker separation in noisy conditions in the time-frequency domain. The inputs to the model include noisy multi-talker mixtures and the corresponding cropped face images. Our approach incorporates cross-attention audiovisual…

Cited by 0SourceScholar
2024

Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark

CVPR 2024highlight

We present a new dataset called Real Acoustic Fields (RAF) that captures real acoustic room data from multiple modalities. The dataset includes high-quality and densely captured room impulse response data paired with multi-view images and precise 6DoF pose tracking data for sound emitters and listen…

Cited by 13SourcePDFScholar
2024

Spherical World-Locking for Audio-Visual Localization in Egocentric Videos

ECCV 2024poster

"Egocentric videos provide comprehensive contexts for user and scene understanding, spanning multisensory perception to behavioral interaction. We propose Spherical World-Locking (SWL) as a general framework for egocentric scene representation, which implicitly transforms multisensory streams with r…

Cited by 4SourcePDFScholar
2023

AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis

NeurIPS 2023poster

Can machines recording an audio-visual scene produce realistic, matching audio-visual experiences at novel positions and novel view directions? We answer it by studying a new task---real-world audio-visual scene synthesis---and a first-of-its-kind NeRF-based approach for multimodal learning. Concret…

Cited by 29SourcePDFScholar
2023

LA-VOCE: LOW-SNR Audio-Visual Speech Enhancement Using Neural Vocoders

ICASSP 2023accepted

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker’s lip movements. This approach has been shown to yield improvements over audio-only speech enhancement, particularly for the removal of interferin…

Cited by 0SourceScholar
2023

Leveraging Heteroscedastic Uncertainty in Learning Complex Spectral Mapping for Single-Channel Speech Enhancement

ICASSP 2023accepted

Most speech enhancement (SE) models learn a point estimate and do not make use of uncertainty estimation in the learning process. In this paper, we show that modeling heteroscedastic uncertainty by minimizing a multivariate Gaussian negative log-likelihood (NLL) improves SE performance at no extra c…

Cited by 2SourceScholar
2023

Nord: Non-Matching Reference Based Relative Depth Estimation from Binaural Speech

ICASSP 2023accepted

We propose NORD: a novel framework for estimating the relative depth between two binaural speech recordings. In contrast to existing depth estimation techniques, ours only requires audio signals as input. We trained the framework to solve depth preference (i.e. which input perceptually sounds closer…

Cited by 0SourceScholar
2023

Paaploss: A Phonetic-Aligned Acoustic Parameter Loss for Speech Enhancement

ICASSP 2023accepted

Despite rapid advancement in recent years, current speech enhancement models often produce speech that differs in perceptual quality from real clean speech. We propose a learning objective that formalizes differences in perceptual quality, by using domain knowledge of acoustic-phonetics. We identify…

Cited by 0SourceScholar
2023

TAPLoss: A Temporal Acoustic Parameter Loss for Speech Enhancement

ICASSP 2023accepted

Speech enhancement models have greatly progressed in recent years, but still show limits in perceptual quality of their speech outputs. We propose an objective for perceptual quality based on temporal acoustic parameters. These are fundamental speech features that play an essential role in various a…

Cited by 0SourceScholar
2023

Torchaudio-Squim: Reference-Less Speech Quality and Intelligibility Measures in Torchaudio

ICASSP 2023accepted

Measuring quality and intelligibility of a speech signal is usually a critical step in development of speech processing systems. To enable this, a variety of metrics to measure quality and intelligibility under different assumptions have been developed. Through this paper, we introduce tools and a s…

Cited by 0SourceScholar
2022

Audio Signal Processing for Telepresence Based on Wearable Array in Noisy and Dynamic Scenes

ICASSP 2022accepted

Telepresence for virtual meetings has gained interest due to recent travel limitations and the new reality of working from home. However, current literature supporting real-world microphone arrays for realistic telepresence in audio is very limited. This paper investigates a scenario of a distant pa…

Cited by 0SourceScholar
2022

Conformer-Based Self-Supervised Learning For Non-Speech Audio Tasks

ICASSP 2022accepted

Representation learning from unlabeled data has been of major interest in artificial intelligence research. While self-supervised speech representation learning has been popular in the speech research community, very few works have comprehensively analyzed audio representation learning for non-speec…

Cited by 0SourceScholar
2022

Continual Self-Training With Bootstrapped Remixing For Speech Enhancement

ICASSP 2022accepted

We propose RemixIT, a simple and novel self-supervised training method for speech enhancement. The proposed method is based on a continuously self-training scheme that overcomes limitations from previous studies including assumptions for the in-domain noise distribution and having access to clean ta…

Cited by 0SourceScholar
2022

Curriculum Optimization for Low-Resource Speech Recognition

ICASSP 2022accepted

Modern end-to-end speech recognition models show astonishing results in transcribing audio signals into written text. However, conventional data feeding pipelines may be sub-optimal for low-resource speech recognition, which still remains a challenging task. We propose an automated curriculum learni…

Cited by 0SourceScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Multichannel Speech Enhancement Without Beamforming

ICASSP 2022accepted

Deep neural networks are often coupled with traditional spatial filters, such as MVDR beamformers for effectively exploiting spatial information. Even though single-stage end-to-end supervised models can obtain impressive enhancement, combining them with a traditional beamformer and a DNN-based post…

Cited by 0SourceScholar
2022

TPARN: Triple-Path Attentive Recurrent Network for Time-Domain Multichannel Speech Enhancement

ICASSP 2022accepted

In this work, we propose a new model called triple-path attentive recurrent network (TPARN) for multichannel speech enhancement in the time domain. TPARN extends a single-channel dual-path network to a multichannel network by adding a third path along the spatial dimension. First, TPARN processes sp…

Cited by 0SourceScholar
2022

The Impact of Removing Head Movements on Audio-Visual Speech Enhancement

ICASSP 2022accepted

This paper investigates the impact of head movements on audio-visual speech enhancement (AVSE). Although being a common conversational feature, head movements have been ignored by past and recent studies: they challenge today’s learning-based methods as they often degrade the performance of models t…

Cited by 0SourceScholar
2021

NORESQA: A Framework for Speech Quality Assessment using Non-Matching References

NeurIPS 2021poster

The perceptual task of speech quality assessment (SQA) is a challenging task for machines to do. Objective SQA methods that rely on the availability of the corresponding clean reference have been the primary go-to approaches for SQA. Clearly, these methods fail in real-world scenarios where the grou…

2020

A Sequential Self Teaching Approach for Improving Generalization in Sound Event Recognition

ICML 2020poster

An important problem in machine auditory perception is to recognize and detect sound events. In this paper, we propose a sequential self-teaching approach to learning sounds. Our main proposition is that it is harder to learn sounds in adverse situations such as from weakly labeled and/or noisy labe…

Cited by 48SourcePDFScholar
2020

SeCoST: : Sequential Co-Supervision for Large Scale Weakly Labeled Audio Event Detection

ICASSP 2020accepted

Weakly supervised learning algorithms are critical for scaling audio event detection to several hundreds of sound categories. Such learning models should not only disambiguate sound events efficiently with minimal class-specific annotation but also be robust to label noise, which is more apparent wi…

Cited by 0SourceScholar
2018

Content-Based Representations of Audio Using Siamese Neural Networks

ICASSP 2018accepted

In this paper, we focus on the problem of content-based retrieval for audio, which aims to retrieve all semantically similar audio recordings for a given audio clip query. This problem is similar to the problem of query by example of audio, which aims to retrieve media samples from a database, which…

Cited by 0SourceScholar
2018

Framework for Evaluation of Sound Event Detection in Web Videos

ICASSP 2018accepted

The largest source of sound events is web videos. Most videos lack sound event labels at segment level, however, a significant number of them do respond to text queries, from a match found using metadata by search engines. In this paper we explore the extent to which a search query can be used as th…

Cited by 0SourceScholar
2018

Knowledge Transfer from Weakly Labeled Audio Using Convolutional Neural Network for Sound Events and Scenes

ICASSP 2018accepted

In this work we propose approaches to effectively transfer knowledge from weakly labeled web audio data. We first describe a convolutional neural network (CNN) based framework for sound event detection and classification using weakly labeled audio data. Our model trains efficiently from audios of va…

Cited by 0SourceScholar