← Search

Vamsi Krishna Ithapu

23 accepted papers

2026

Forecasting 3D Scanpaths in Egocentric Video

CVPR 2026

Forecasting gaze behavior is an important task for understanding user intent and creating AR/VR systems that can anticipate where users will look and interact next. While prior works have addressed predicting scanpaths in static images, forecasting gaze in egocentric videos presents new challenges d

Cited by 0SourcecodeScholar
2026

MORE THAN A SHORTCUT: A HYPERBOLIC APPROACH TO EARLY-EXIT NETWORKS

ICASSP 2026poster

Deploying accurate event detection on resource-constrained devices is challenged by the trade-off between performance and computational cost. While Early-Exit (EE) networks offer a solution through adaptive computation, they often fail to enforce a coherent hierarchical structure, limiting the relia…

Cited by 0SourcePDFScholar
2025

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

ICCV 2025poster

Modern perception models, particularly those designed for multisensory egocentric tasks, have achieved remarkable performance but often come with substantial computational costs. These high demands pose challenges for real-world deployment, especially in resource-constrained environments. In this pa…

Cited by 0SourcePDFScholar
2025

Hearing Anywhere in Any Environment

CVPR 2025poster

In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR) estimation, most existing methods are limited to the single environ…

Cited by 0SourcePDFScholar
2025

Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech Enhancement

ICASSP 2025accepted

Deep learning-based speech enhancement (SE) methods often face significant computational challenges when needing to meet low-latency requirements because of the increased number of frames to be processed. This paper introduces the SlowFast framework which aims to reduce computation costs specificall…

Cited by 0SourceScholar
2024

Hearing Loss Detection From Facial Expressions in One-On-One Conversations

ICASSP 2024accepted

Individuals with impaired hearing experience difficulty in conversations, especially in noisy environments. This difficulty often manifests as a change in behavior and may be captured via facial expressions, such as the expression of discomfort or fatigue. In this work, we build on this idea and int…

Cited by 0SourceScholar
2024

Self-Motion As Supervision For Egocentric Audiovisual Localization

ICASSP 2024accepted

Sound source localization is a key requirement for many assistive applications of augmented reality, such as speech enhancement. In conversational settings, potential sources of interest may be approximated by active speaker detection. However, localizing speakers in crowded, noisy environments is c…

Cited by 0SourceScholar
2024

Spherical World-Locking for Audio-Visual Localization in Egocentric Videos

ECCV 2024poster

"Egocentric videos provide comprehensive contexts for user and scene understanding, spanning multisensory perception to behavioral interaction. We propose Spherical World-Locking (SWL) as a general framework for egocentric scene representation, which implicitly transforms multisensory streams with r…

Cited by 4SourcePDFScholar
2024

The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective

CVPR 2024poster

In recent years the thriving development of research related to egocentric videos has provided a unique perspective for the study of conversational interactions where both visual and audio signals play a crucial role. While most prior work focus on learning about behaviors that directly involve the…

2023

Chat2Map: Efficient Scene Mapping From Multi-Ego Conversations

CVPR 2023poster

Can conversational videos captured from multiple egocentric viewpoints reveal the map of a scene in a cost-efficient way? We seek to answer this question by proposing a new problem: efficiently building the map of a previously unseen 3D environment by exploiting shared information in the egocentric…

Cited by 9SourcePDFScholar
2023

Egocentric Auditory Attention Localization in Conversations

CVPR 2023poster

In a noisy conversation environment such as a dinner party, people often exhibit selective auditory attention, or the ability to focus on a particular speaker while tuning out others. Recognizing who somebody is listening to in a conversation is essential for developing technologies that can underst…

2023

LA-VOCE: LOW-SNR Audio-Visual Speech Enhancement Using Neural Vocoders

ICASSP 2023accepted

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker’s lip movements. This approach has been shown to yield improvements over audio-only speech enhancement, particularly for the removal of interferin…

Cited by 0SourceScholar
2023

Leveraging Heteroscedastic Uncertainty in Learning Complex Spectral Mapping for Single-Channel Speech Enhancement

ICASSP 2023accepted

Most speech enhancement (SE) models learn a point estimate and do not make use of uncertainty estimation in the learning process. In this paper, we show that modeling heteroscedastic uncertainty by minimizing a multivariate Gaussian negative log-likelihood (NLL) improves SE performance at no extra c…

Cited by 0SourceScholar
2023

Novel-View Acoustic Synthesis

CVPR 2023poster

We introduce the novel-view acoustic synthesis (NVAS) task: given the sight and sound observed at a source viewpoint, can we synthesize the sound of that scene from an unseen target viewpoint? We propose a neural rendering approach: Visually-Guided Acoustic Synthesis (ViGAS) network that learns to s…

2023

Towards Improved Room Impulse Response Estimation for Speech Recognition

ICASSP 2023accepted

We propose a novel approach for blind room impulse response (RIR) estimation systems in the context of a downstream application scenario, far-field automatic speech recognition (ASR). We first draw the connection between improved RIR estimation and improved ASR performance, as a means of evaluating…

Cited by 0SourceScholar
2022

Deep Impulse Responses: Estimating and Parameterizing Filters with Deep Networks

ICASSP 2022accepted

Impulse response estimation in high noise and in-the-wild settings, with minimal control of the underlying data distributions, is a challenging problem. We propose a novel framework for parameterizing and estimating impulse responses based on recent advances in neural representation learning. Our fr…

Cited by 0SourceScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2021

Audio-Visual Floorplan Reconstruction

ICCV 2021poster

Given only a few glimpses of an environment, how much can we infer about its entire floorplan? Existing methods can map only what is visible or immediately apparent from context, and thus require substantial movements through a space to fully map it. We explore how both audio and visual sensing toge…

Cited by 57PDFScholar
2020

SeCoST: : Sequential Co-Supervision for Large Scale Weakly Labeled Audio Event Detection

ICASSP 2020accepted

Weakly supervised learning algorithms are critical for scaling audio event detection to several hundreds of sound categories. Such learning models should not only disambiguate sound events efficiently with minimal class-specific annotation but also be robust to label noise, which is more apparent wi…

Cited by 0SourceScholar
2020

SoundSpaces: Audio-Visual Navigation in 3D Environments

ECCV 2020poster

Moving around in the world is naturally a multi-sensory experience, but today's embodied agents are deaf - restricted to solely their visual perception of the environment. We introduce audio-visual navigation for complex, acoustically and visually realistic 3D environments. By both seeing and hearin…