← Search

Kristen Grauman

110 accepted papers

2026

Mash, Spread, Slice! Learning to Manipulate Object States Via Visual Spatial Progress

ICRA 2026poster

Most robot manipulation focuses on changing the kinematic state of objects: picking, placing, opening, or rotating them. However, a wide range of real-world manipulation tasks involve a different class of object state change—such as mashing, spreading, or slicing—where the object’s physical and visu…

2025

ExpertAF: Expert Actionable Feedback from Video

CVPR 2025poster

Feedback is essential for learning a new skill or improving one's current skill-level. However, current methods for skill-assessment from video only provide scores or compare demonstrations, leaving the burden of knowing what to do differently on the user. We introduce a novel method to generate act…

Cited by 3SourcePDFScholar
2025

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

NeurIPS 2025spotlight

Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The research community has responded by using distillation from black-box models to label training data, achieving strong benchmark…

Cited by 0SourcecodeScholar
2025

Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos

ICCV 2025poster

We introduce Switch-a-View, a model that learns to automatically select the viewpoint to display at each timepoint when creating a how-to video. The key insight of our approach is how to train such a model from unlabeled--but human-edited--video samples. We pose a pretext task that pseudo-labels seg…

Cited by 0SourcePDFScholar
2025

Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation Learning

CVPR 2025poster

Egocentric and exocentric perspectives of human action differ significantly, yet overcoming this extreme viewpoint gap is critical for applications in augmented reality and robotics. We propose ViewpointRosetta, an approach that unlocks large-scale unpaired ego and exo video data to learn clip-level…

Cited by 0SourcePDFScholar
2025

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning

NeurIPS 2025poster

Video reasoning, the task of enabling machines to infer from dynamic visual content through multi-step logic, is crucial for advanced AI. While the Chain-of-Thought (CoT) mechanism has enhanced reasoning in text-based tasks, its application to video understanding remains underexplored. This paper pr…

Cited by 0SourceScholar
2025

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

CVPR 2025highlight

Given a multi-view video, which viewpoint is most informative for a human observer? Existing methods rely on heuristics or expensive "best-view" supervision to answer this question, limiting their applicability. We propose a weakly supervised approach that leverages language accompanying an instruct…

Cited by 0SourcePDFScholar
2024

4Diff: 3D-Aware Diffusion Model for Third-to-First Viewpoint Translation

ECCV 2024poster

"We present , a 3D-aware diffusion model addressing the exo-to-ego viewpoint translation task — generating first-person (egocentric) view images from the corresponding third-person (exocentric) images. Building on the diffusion model’s ability to generate photorealistic images, we propose a transfor…

2024

Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos

ECCV 2024oral

"Generating realistic audio for human actions is important for many applications, such as creating sound effects for films or virtual reality games. Existing approaches implicitly assume total correspondence between the video and audio during training, yet many sounds happen off-screen and have weak…

2024

ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling

IROS 2024

An environment acoustic model represents how sound is transformed by the physical characteristics of an indoor environment, for any given source/receiver location. Traditional methods for constructing acoustic models involve expensive and time-consuming collection of large quantities of acoustic dat

Cited by 2SourceScholar
2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness

NeurIPS 2024poster

We study the problem of precisely swapping objects in videos, with a focus on those interacted with by hands, given one user-provided reference object image. Despite the great advancements that diffusion models have made in video editing recently, these models often fall short in handling the intric…

Cited by 8SourcePDFScholar
2024

Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos

CVPR 2024poster

We propose a self-supervised method for learning representations based on spatial audio-visual correspondences in egocentric videos. Our method uses a masked auto-encoding framework to synthesize masked binaural audio through the synergy of audio and vision thereby learning useful spatial relationsh…

Cited by 6SourcePDFScholar
2024

Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos

ECCV 2024poster

"We investigate exocentric-to-egocentric cross-view translation, which aims to generate a first-person (egocentric) view of an actor based on a video recording that captures the actor from a third-person (exocentric) perspective. To this end, we propose a generative framework called Exo2Ego that dec…

Cited by 18SourcePDFScholar
2024

Sim2Real Transfer for Audio-Visual Navigation with Frequency-Adaptive Acoustic Field Prediction

IROS 2024poster

Sim2real transfer has received increasing attention lately due to its success in transferring robotic policies learned in simulation to the real world. While significant progress has been made in transferring vision-based navigation policies, the current sim2real strategy for audio-visual navigation…

Cited by 4SourceScholar
2024

SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos

CVPR 2024poster

We propose a novel self-supervised embedding to learn how actions sound from narrated in-the-wild egocentric videos. Whereas existing methods rely on curated data with known audio-visual correspondence our multimodal contrastive-consensus coding (MC3) embedding reinforces the associations between au…

Cited by 8SourcePDFScholar
2023

Chat2Map: Efficient Scene Mapping From Multi-Ego Conversations

CVPR 2023poster

Can conversational videos captured from multiple egocentric viewpoints reveal the map of a scene in a cost-efficient way? We seek to answer this question by proposing a new problem: efficiently building the map of a previously unseen 3D environment by exploiting shared information in the egocentric…

Cited by 9SourcePDFScholar
2023

EgoDistill: Egocentric Head Motion Distillation for Efficient Video Understanding

NeurIPS 2023poster

Recent advances in egocentric video understanding models are promising, but their heavy computational expense is a barrier for many real-world applications. To address this challenge, we propose EgoDistill, a distillation-based approach that learns to reconstruct heavy ego-centric video clip feature…

Cited by 25SourcePDFScholar
2023

EgoEnv: Human-centric environment representations from egocentric video

NeurIPS 2023oral

First-person video highlights a camera-wearer's activities in the context of their persistent environment. However, current video understanding approaches reason over visual features from short video clips that are detached from the underlying physical space and capture only what is immediately vis…

Cited by 20SourcePDFScholar
2023

EgoTracks: A Long-term Egocentric Visual Object Tracking Dataset

NeurIPS 2023poster

Visual object tracking is a key component to many egocentric vision problems. However, the full spectrum of challenges of egocentric tracking faced by an embodied AI is underrepresented in many existing datasets; these tend to focus on relatively short, third-person videos. Egocentric video has seve…

2023

HierVL: Learning Hierarchical Video-Language Embeddings

CVPR 2023highlight

Video-language embeddings are a promising avenue for injecting semantics into visual representations, but existing methods capture only short-term associations between seconds-long video clips and their accompanying text. We propose HierVL, a novel hierarchical video-language embedding that simultan…

Cited by 57SourcePDFScholar
2023

Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Alignment

NeurIPS 2023poster

The egocentric and exocentric viewpoints of a human activity look dramatically different, yet invariant representations to link them are essential for many potential applications in robotics and augmented reality. Prior work is limited to learning view-invariant features from paired synchronized vi…

Cited by 37SourcePDFScholar
2023

NaQ: Leveraging Narrations As Queries To Supervise Episodic Memory

CVPR 2023poster

Searching long egocentric videos with natural language queries (NLQ) has compelling applications in augmented reality and robotics, where a fluid index into everything that a person (agent) has seen before could augment human memory and surface relevant information on demand. However, the structured…

2023

Novel-View Acoustic Synthesis

CVPR 2023poster

We introduce the novel-view acoustic synthesis (NVAS) task: given the sight and sound observed at a source viewpoint, can we synthesize the sound of that scene from an unseen target viewpoint? We propose a neural rendering approach: Visually-Guided Acoustic Synthesis (ViGAS) network that learns to s…

2023

Single-Stage Visual Query Localization in Egocentric Videos

NeurIPS 2023poster

Visual Query Localization on long-form egocentric videos requires spatio-temporal search and localization of visually specified objects and is vital to build episodic memory systems. Prior work develops complex multi-stage pipelines that leverage well-established object detection and tracking method…

Cited by 17SourcePDFScholar
2023

SpotEM: Efficient Video Search for Episodic Memory

ICML 2023poster

The goal in episodic memory (EM) is to search a long egocentric video to answer a natural language query (e.g., “where did I leave my purse?”). Existing EM methods exhaustively extract expensive fixed-length clip features to look everywhere in the video for the answer, which is infeasible for long w…

Cited by 11SourcePDFScholar
2023

Video-Mined Task Graphs for Keystep Recognition in Instructional Videos

NeurIPS 2023poster

Procedural activity understanding requires perceiving human actions in terms of a broader task, where multiple keysteps are performed in sequence across a long video to reach a final goal state---such as the steps of a recipe or the steps of a DIY fix-it task. Prior work largely treats keystep reco…

Cited by 29SourcePDFScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Egocentric Activity Recognition and Localization on a 3D Map

ECCV 2022poster

"Given a video captured from a first person perspective and the environment context of where the video is recorded, can we recognize what the person is doing and identify where the action occurs in the 3D space? We address this challenging problem of jointly recognizing and localizing actions of a m…

Cited by 27SourcePDFScholar
2022

Environment Predictive Coding for Visual Navigation

ICLR 2022poster

We introduce environment predictive coding, a self-supervised approach to learn environment-level representations for embodied agents. In contrast to prior work on self-supervised learning for individual images, we aim to encode a 3D environment using a series of images observed by an agent moving i…

Cited by 8SourcePDFScholar
2022

Few-Shot Audio-Visual Learning of Environment Acoustics

NeurIPS 2022accept

Room impulse response (RIR) functions capture how the surrounding physical environment transforms the sounds heard by a listener, with implications for various applications in AR, VR, and robotics. Whereas traditional methods to estimate RIRs assume dense geometry and/or sound measurements throughou…

Cited by 55SourcePDFScholar
2022

PONI: Potential Functions for ObjectGoal Navigation With Interaction-Free Learning

CVPR 2022oral

State-of-the-art approaches to ObjectGoal navigation (ObjectNav) rely on reinforcement learning and typically require significant computational resources and time for learning. We propose Potential functions for ObjectGoal Navigation with Interaction-free learning (PONI), a modular approach that dis…

Cited by 178PDFcodeScholar
2022

SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning

NeurIPS 2022accept

We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from arbitrary microphone locations. Together with existing 3D vi…

Cited by 94SourcePDFScholar
2022

Zero Experience Required: Plug & Play Modular Transfer Learning for Semantic Visual Navigation

CVPR 2022poster

In reinforcement learning for visual navigation, it is common to develop a model for each new task, and train that model from scratch with task-specific interactions in 3D environments. However, this process is expensive; massive amounts of interactions are needed for the model to generalize well. M…

Cited by 84PDFScholar
2021

Audio-Visual Floorplan Reconstruction

ICCV 2021poster

Given only a few glimpses of an environment, how much can we infer about its entire floorplan? Existing methods can map only what is visible or immediately apparent from context, and thus require substantial movements through a space to fully map it. We explore how both audio and visual sensing toge…

Cited by 57PDFScholar
2021

Ego-Exo: Transferring Visual Representations From Third-Person to First-Person Videos

CVPR 2021poster

We introduce an approach for pre-training egocentric video models using large-scale third-person video datasets. Learning from purely egocentric data is limited by low dataset scale and diversity, while using purely exocentric (third-person) data introduces a large domain mismatch. Our idea is to di…

Cited by 98PDFcodeScholar
2021

Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback

CVPR 2021poster

Conversational interfaces for the detail-oriented retail fashion domain are more natural, expressive, and user friendly than classical keyword-based search interfaces. In this paper, we introduce the Fashion IQ dataset to support and advance research on interactive fashion image retrieval. Fashion I…

Cited by 297PDFcodeScholar
2021

Learning to Set Waypoints for Audio-Visual Navigation

ICLR 2021poster

In audio-visual navigation, an agent intelligently travels through a complex, unmapped 3D environment using both sights and sounds to find a sound source (e.g., a phone ringing in another room). Existing models learn to act at a fixed granularity of agent motion and rely on simple recurrent aggregat…

2021

Multiview Pseudo-Labeling for Semi-Supervised Learning From Video

ICCV 2021poster

We present a multiview pseudo-labeling approach to video learning, a novel framework that uses complementary views in the form of appearance and motion information for semi-supervised learning in video. The complementary views help obtain more reliable "pseudo-labels"" on unlabeled video, to learn s…

Cited by 63PDFScholar
2021

Shaping embodied agent behavior with activity-context priors from egocentric video

NeurIPS 2021spotlight

Complex physical tasks entail a sequence of object interactions, each with its own preconditions -- which can be difficult for robotic agents to learn efficiently solely through their own experience. We introduce an approach to discover activity-context priors from in-the-wild egocentric video captu…

Cited by 18SourcePDFScholar
2020

Don't Judge an Object by Its Context: Learning to Overcome Contextual Bias

CVPR 2020oral

Existing models often leverage co-occurrences between objects and their context to improve recognition accuracy. However, strongly relying on context risks a model's generalizability, especially when typical co-occurrence patterns are absent. This work focuses on addressing such contextual biases to…

Cited by 141PDFScholar
2020

Ego-Topo: Environment Affordances From Egocentric Video

CVPR 2020oral

First-person video naturally brings the use of a physical environment to the forefront, since it shows the camera wearer interacting fluidly in a space based on his intentions. However, current methods largely separate the observed actions from the persistent space itself. We introduce a model for e…

Cited by 152PDFcodeScholar
2020

Learning Affordance Landscapes for Interaction Exploration in 3D Environments

NeurIPS 2020spotlight

Embodied agents operating in human spaces must be able to master how their environment works: what objects can the agent use, and how can it use them? We introduce a reinforcement learning approach for exploration for interaction, whereby an embodied agent autonomously discovers the affordance lands…

Cited by 92SourcePDFScholar
2020

Occupancy Anticipation for Efficient Exploration and Navigation

ECCV 2020poster

State-of-the-art navigation methods leverage a spatial memory to generalize to new environments, but their occupancy maps are limited to capturing the geometric structures directly observed by the agent. We propose occupancy anticipation, where the agent uses its egocentric RGB-D observations to inf…

2020

VisualEchoes: Spatial Image Representation Learning through Echolocation

ECCV 2020poster

Several animal species (e.g., bats, dolphins, and whales) and even visually impaired humans have the remarkable ability to perform echolocation: a biological sonar used to perceive spatial layout and locate objects in the world. We explore the spatial cues contained in echoes and how they can benefi…

2020

You2Me: Inferring Body Pose in Egocentric Video via First and Second Person Interactions

CVPR 2020oral

The body pose of a person wearing a camera is of great interest for applications in augmented reality, healthcare, and robotics, yet much of the person's body is out of view for a typical wearable camera. We propose a learning-based approach to estimate the camera wearer's 3D body pose from egocentr…

Cited by 113PDFcodeScholar
2019

2.5D Visual Sound

CVPR 2019oral

Binaural audio provides a listener with 3D sound sensation, allowing a rich perceptual experience of the scene. However, binaural recordings are scarcely available and require nontrivial expertise and equipment to obtain. We propose to convert common monaural audio into binaural audio by leveraging…

Cited by 261PDFcodeScholar
2019

Extreme Relative Pose Estimation for RGB-D Scans via Scene Completion

CVPR 2019oral

Estimating the relative rigid pose between two RGB-D scans of the same underlying environment is a fundamental problem in computer vision, robotics, and computer graphics. Most existing approaches allow only limited maximum relative pose changes since they require considerable overlap between the in…

Cited by 57PDFcodeScholar
2019

Less Is More: Learning Highlight Detection From Video Duration

CVPR 2019poster

Highlight detection has the potential to significantly ease video browsing, but existing methods often suffer from expensive supervision requirements, where human viewers must manually identify highlights in training videos. We propose a scalable unsupervised solution that exploits video duration as…

Cited by 159PDFScholar
2019

SpotTune: Transfer Learning Through Adaptive Fine-Tuning

CVPR 2019poster

Transfer learning, which allows a source task to affect the inductive bias of the target task, is widely used in computer vision. The typical way of conducting transfer learning with deep neural networks is to fine-tune a model pretrained on the source task using data from the target task. In this p…

Cited by 640PDFScholar
2018

Attributes as Operators: Factorizing Unseen Attribute-Object Compositions

ECCV 2018poster

We present a new approach to modeling visual attributes. Prior work casts attributes in a similar role as objects, learning a latent representation where properties (e.g., sliced) are recognized by classifiers much in the way objects (e.g., apple) are. However, this common approach fails to separate…

2018

BlockDrop: Dynamic Inference Paths in Residual Networks

CVPR 2018poster

Very deep convolutional neural networks offer excellent recognition results, yet their computational expense limits their impact for many real-world applications. We introduce BlockDrop, an approach that learns to dynamically choose which layers of a deep network to execute during inference so as t…

2018

Im2Flow: Motion Hallucination From Static Images for Action Recognition

CVPR 2018poster

Existing methods to recognize actions in static images take the images at their face value, learning the appearances---objects, scenes, and body poses---that distinguish each action class. However, such models are deprived of the rich dynamic structure and motions that also define human activity. We…

2018

Learning to Look Around: Intelligently Exploring Unseen Environments for Unknown Tasks

CVPR 2018poster

It is common to implicitly assume access to intelligently captured inputs (e.g., photos from a human photographer), yet autonomously capturing good observations is itself a major challenge. We address the problem of learning to look around: if an agent has the ability to voluntarily acquire new vie…

Cited by 130SourcePDFScholar
2018

Learning to Separate Object Sounds by Watching Unlabeled Video

ECCV 2018poster

Perceiving a scene most fully requires all the senses. Yet modeling how objects look and sound is challenging: most natural scenes and events contain multiple objects, and the audio track mixes all the sound sources together. We propose to learn audio-visual object models from unlabeled video, then…

2018

ShapeCodes: Self-Supervised Feature Learning by Lifting Views to Viewgrids

ECCV 2018poster

We introduce an unsupervised feature learning approach that embeds 3D shape information into a single-view image representation. The main idea is a self-supervised training objective that, given only a single 2D image, requires all unseen views of the object to be predictable from learned features.…

Cited by 20SourcePDFScholar
2018

VizWiz Grand Challenge: Answering Visual Questions From Blind People

CVPR 2018poster

The study of algorithms to automatically answer visual questions currently is motivated by visual question answering (VQA) datasets constructed in artificial VQA settings. We propose VizWiz, the first goal-oriented VQA dataset arising from a natural VQA setting. VizWiz consists of 31,000 visual qu…

Cited by 961SourcePDFScholar
2017

FusionSeg: Learning to Combine Motion and Appearance for Fully Automatic Segmentation of Generic Objects in Videos

CVPR 2017poster

We propose an end-to-end learning framework for segmenting generic objects in videos. Our method learns to combine appearance and motion information to produce pixel level segmentation masks for all prominent objects in videos. We formulate this task as a structured prediction problem and design a t…

Cited by 477PDFScholar
2017

Learning Spherical Convolution for Fast Features from 360° Imagery

NeurIPS 2017poster

While 360° cameras offer tremendous new possibilities in vision, graphics, and augmented reality, the spherical images they produce make core feature extraction non-trivial. Convolutional neural networks (CNNs) trained on images from perspective cameras yield “flat" filters, yet 360° images cannot b…

2017

Learning the Latent "Look": Unsupervised Discovery of a Style-Coherent Embedding From Fashion Images

ICCV 2017poster

What defines a visual style? Fashion styles emerge organically from how people assemble outfits of clothing, making them difficult to pin down with a computational model. Low-level visual similarity can be too specific to detect stylistically similar images, while manually crafted style categories c…

Cited by 114PDFScholar
2016

Pull the Plug? Predicting If Computers or Humans Should Segment Images

CVPR 2016poster

Foreground object segmentation is a critical step for many image analysis tasks. While automated methods can produce high-quality results, their failures disappoint users in need of practical solutions. We propose a resource allocation framework for predicting how best to allocate a fixed budget o…

Cited by 34PDFScholar
2016

Summary Transfer: Exemplar-Based Subset Selection for Video Summarization

CVPR 2016poster

Video summarization has unprecedented importance to help us digest, browse, and search today's ever-growing video collections. We propose a novel subset selection technique that leverages supervision in the form of human-created summaries to perform automatic keyframe-based video summarization. The…

Cited by 271PDFScholar