← Search

Ziad Al-Halah

22 accepted papers

2025

How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes

ICCV 2025poster

How would the sound in a studio change with a carpeted floor and acoustic tiles on the walls? We introduce the task of material-controlled acoustic profile generation, where, given an indoor scene with specific audio-visual characteristics, the goal is to generate a target acoustic profile based on…

2025

Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos

ICCV 2025poster

We introduce Switch-a-View, a model that learns to automatically select the viewpoint to display at each timepoint when creating a how-to video. The key insight of our approach is how to train such a model from unlabeled--but human-edited--video samples. We pose a pretext task that pseudo-labels seg…

Cited by 0SourcePDFScholar
2025

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

CVPR 2025highlight

Given a multi-view video, which viewpoint is most informative for a human observer? Existing methods rely on heuristics or expensive "best-view" supervision to answer this question, limiting their applicability. We propose a weakly supervised approach that leverages language accompanying an instruct…

Cited by 0SourcePDFScholar
2024

Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos

CVPR 2024poster

We propose a self-supervised method for learning representations based on spatial audio-visual correspondences in egocentric videos. Our method uses a masked auto-encoding framework to synthesize masked binaural audio through the synergy of audio and vision thereby learning useful spatial relationsh…

Cited by 6SourcePDFScholar
2023

NaQ: Leveraging Narrations As Queries To Supervise Episodic Memory

CVPR 2023poster

Searching long egocentric videos with natural language queries (NLQ) has compelling applications in augmented reality and robotics, where a fluid index into everything that a person (agent) has seen before could augment human memory and surface relevant information on demand. However, the structured…

2023

SpotEM: Efficient Video Search for Episodic Memory

ICML 2023poster

The goal in episodic memory (EM) is to search a long egocentric video to answer a natural language query (e.g., “where did I leave my purse?”). Existing EM methods exhaustively extract expensive fixed-length clip features to look everywhere in the video for the answer, which is infeasible for long w…

Cited by 11SourcePDFScholar
2022

Environment Predictive Coding for Visual Navigation

ICLR 2022poster

We introduce environment predictive coding, a self-supervised approach to learn environment-level representations for embodied agents. In contrast to prior work on self-supervised learning for individual images, we aim to encode a 3D environment using a series of images observed by an agent moving i…

Cited by 8SourcePDFScholar
2022

Few-Shot Audio-Visual Learning of Environment Acoustics

NeurIPS 2022accept

Room impulse response (RIR) functions capture how the surrounding physical environment transforms the sounds heard by a listener, with implications for various applications in AR, VR, and robotics. Whereas traditional methods to estimate RIRs assume dense geometry and/or sound measurements throughou…

Cited by 55SourcePDFScholar
2022

PONI: Potential Functions for ObjectGoal Navigation With Interaction-Free Learning

CVPR 2022oral

State-of-the-art approaches to ObjectGoal navigation (ObjectNav) rely on reinforcement learning and typically require significant computational resources and time for learning. We propose Potential functions for ObjectGoal Navigation with Interaction-free learning (PONI), a modular approach that dis…

Cited by 178PDFcodeScholar
2022

Zero Experience Required: Plug & Play Modular Transfer Learning for Semantic Visual Navigation

CVPR 2022poster

In reinforcement learning for visual navigation, it is common to develop a model for each new task, and train that model from scratch with task-specific interactions in 3D environments. However, this process is expensive; massive amounts of interactions are needed for the model to generalize well. M…

Cited by 84PDFScholar
2021

Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback

CVPR 2021poster

Conversational interfaces for the detail-oriented retail fashion domain are more natural, expressive, and user friendly than classical keyword-based search interfaces. In this paper, we introduce the Fashion IQ dataset to support and advance research on interactive fashion image retrieval. Fashion I…

Cited by 297PDFcodeScholar
2021

Learning to Set Waypoints for Audio-Visual Navigation

ICLR 2021poster

In audio-visual navigation, an agent intelligently travels through a complex, unmapped 3D environment using both sights and sounds to find a sound source (e.g., a phone ringing in another room). Existing models learn to act at a fixed granularity of agent motion and rely on simple recurrent aggregat…

2020

Occupancy Anticipation for Efficient Exploration and Navigation

ECCV 2020poster

State-of-the-art navigation methods leverage a spatial memory to generalize to new environments, but their occupancy maps are limited to capturing the geometric structures directly observed by the agent. We propose occupancy anticipation, where the agent uses its egocentric RGB-D observations to inf…

2020

SoundSpaces: Audio-Visual Navigation in 3D Environments

ECCV 2020poster

Moving around in the world is naturally a multi-sensory experience, but today's embodied agents are deaf - restricted to solely their visual perception of the environment. We introduce audio-visual navigation for complex, acoustically and visually realistic 3D environments. By both seeing and hearin…

2020

VisualEchoes: Spatial Image Representation Learning through Echolocation

ECCV 2020poster

Several animal species (e.g., bats, dolphins, and whales) and even visually impaired humans have the remarkable ability to perform echolocation: a biological sonar used to perceive spatial layout and locate objects in the world. We explore the spatial cues contained in echoes and how they can benefi…

2017

Automatic Discovery, Association Estimation and Learning of Semantic Attributes for a Thousand Categories

CVPR 2017poster

Attribute-based recognition models, due to their impressive performance and their ability to generalize well on novel categories, have been widely adopted for many computer vision applications. However, usually both the attribute vocabulary and the class-attribute associations have to be provided ma…

Cited by 38PDFScholar
2016

Recovering the Missing Link: Predicting Class-Attribute Associations for Unsupervised Zero-Shot Learning

CVPR 2016accepted

Collecting training images for all visual categories is not only expensive but also impractical. Zero-shot learning (ZSL), especially using attributes, offers a pragmatic solution to this problem. However, at test time most attribute-based methods require a full description of attribute associations…

Cited by 126SourcePDFScholar