← Search

Ishwarya Ananthabhotla

11 accepted papers

2026

Forecasting 3D Scanpaths in Egocentric Video

CVPR 2026

Forecasting gaze behavior is an important task for understanding user intent and creating AR/VR systems that can anticipate where users will look and interact next. While prior works have addressed predicting scanpaths in static images, forecasting gaze in egocentric videos presents new challenges d

Cited by 0SourcecodeScholar
2025

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

ICCV 2025poster

Modern perception models, particularly those designed for multisensory egocentric tasks, have achieved remarkable performance but often come with substantial computational costs. These high demands pose challenges for real-world deployment, especially in resource-constrained environments. In this pa…

Cited by 0SourcePDFScholar
2025

Hearing Anywhere in Any Environment

CVPR 2025poster

In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR) estimation, most existing methods are limited to the single environ…

Cited by 0SourcePDFScholar
2025

SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding

CVPR 2025highlight

We introduce SoundVista, a method to generate the ambient sound of an arbitrary scene at novel viewpoints. Given a pre-acquired recording of the scene from sparsely distributed microphones, SoundVista can synthesize the sound of that scene from an unseen target viewpoint. The method learns the under…

Cited by 0SourcePDFScholar
2024

Hearing Loss Detection From Facial Expressions in One-On-One Conversations

ICASSP 2024accepted

Individuals with impaired hearing experience difficulty in conversations, especially in noisy environments. This difficulty often manifests as a change in behavior and may be captured via facial expressions, such as the expression of discomfort or fatigue. In this work, we build on this idea and int…

Cited by 0SourceScholar
2024

On HRTF Notch Frequency Prediction using Anthropometric Features and Neural Networks

ICASSP 2024accepted

High fidelity spatial audio often performs better when produced using a personalized head-related transfer function (HRTF). However, the direct acquisition of HRTFs is cumbersome and requires specialized equipment. Thus, many personalization methods estimate HRTF features from easily obtained anthro…

Cited by 0SourceScholar
2024

Self-Motion As Supervision For Egocentric Audiovisual Localization

ICASSP 2024accepted

Sound source localization is a key requirement for many assistive applications of augmented reality, such as speech enhancement. In conversational settings, potential sources of interest may be approximated by active speaker detection. However, localizing speakers in crowded, noisy environments is c…

Cited by 0SourceScholar
2024

Spherical World-Locking for Audio-Visual Localization in Egocentric Videos

ECCV 2024poster

"Egocentric videos provide comprehensive contexts for user and scene understanding, spanning multisensory perception to behavioral interaction. We propose Spherical World-Locking (SWL) as a general framework for egocentric scene representation, which implicitly transforms multisensory streams with r…

Cited by 4SourcePDFScholar
2024

The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective

CVPR 2024poster

In recent years the thriving development of research related to egocentric videos has provided a unique perspective for the study of conversational interactions where both visual and audio signals play a crucial role. While most prior work focus on learning about behaviors that directly involve the…

2023

Towards Improved Room Impulse Response Estimation for Speech Recognition

ICASSP 2023accepted

We propose a novel approach for blind room impulse response (RIR) estimation systems in the context of a downstream application scenario, far-field automatic speech recognition (ASR). We first draw the connection between improved RIR estimation and improved ASR performance, as a means of evaluating…

Cited by 0SourceScholar
2019

HCU400: an Annotated Dataset for Exploring Aural Phenomenology through Causal Uncertainty

ICASSP 2019accepted

The way we perceive a sound depends on many aspects- its ecological frequency, acoustic features, typicality, and most notably, its identified source. In this paper, we present the HCU400: a dataset of 402 sounds ranging from easily identifiable everyday sounds to intentionally obscured artificial o…

Cited by 0SourceScholar