← Search

Adam W. Harley

20 accepted papers

2025

AllTracker: Efficient Dense Point Tracking at High Resolution

ICCV 2025poster

We introduce AllTracker: a model that estimates long-range point tracks by way of estimating the flow field between a query frame and every other frame of a video. Unlike existing point tracking methods, our approach delivers high-resolution and dense (all-pixel) correspondence fields, which can be…

2025

LookOut: Real-World Humanoid Egocentric Navigation

ICCV 2025poster

The ability to predict collision-free future trajectories from egocentric observations is crucial in applications such as humanoid robotics, VR / AR, and assistive navigation. In this work, we introduce the challenging problem of predicting a sequence of future 6D head poses from an egocentric video…

2025

MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds

CVPR 2025highlight

We introduce 4D Motion Scaffolds (MoSca), a modern 4D reconstruction system designed to reconstruct and synthesize novel views of dynamic scenes from monocular videos captured casually in the wild. To address such a challenging and ill-posed inverse problem, we leverage prior knowledge from foundati…

2025

TAPIP3D: Tracking Any Point in Persistent 3D Geometry

NeurIPS 2025poster

We introduce TAPIP3D, a novel approach for long-term 3D point tracking in monocular RGB and RGB-D videos. TAPIP3D represents videos as camera-stabilized spatio-temporal feature clouds, leveraging depth and camera motion information to lift 2D video features into a 3D world space where camera moveme…

Cited by 0SourcecodeScholar
2024

ODIN: A Single Model for 2D and 3D Segmentation

CVPR 2024highlight

State-of-the-art models on contemporary 3D segmentation benchmarks like ScanNet consume and label dataset-provided 3D point clouds obtained through post processing of sensed multiview RGB-D images. They are typically trained in-domain forego large-scale 2D pre-training and outperform alternatives th…

2024

Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models

ICRA 2024poster

Object tracking is central to robot perception and scene understanding, allowing robots to parse a video stream in terms of moving objects with names. Tracking-by-detection has long been a dominant paradigm for object tracking of specific object categories [1], [2]. Recently, large-scale pre-trained…

Cited by 12SourceScholar
2023

PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point Tracking

ICCV 2023oral

We introduce PointOdyssey, a large-scale synthetic dataset, and data generation framework, for the training and evaluation of long-term fine-grained tracking algorithms. Our goal is to advance the state-of-the-art by placing emphasis on long videos with naturalistic motion. Toward the goal of natura…

Cited by 143PDFcodeScholar
2023

Simple-BEV: What Really Matters for Multi-Sensor BEV Perception?

ICRA 2023poster

Building 3D perception systems for autonomous vehicles that do not rely on high-density LiDAR is a critical research problem because of the expense of LiDAR systems compared to cameras and other sensors. Recent research has developed a variety of camera-only methods, where features are differentiabl…

Cited by 143SourceScholar
2022

Particle Video Revisited: Tracking through Occlusions Using Point Trajectories

ECCV 2022poster

"Tracking pixels in videos is typically studied as an optical flow estimation problem, where every pixel is described with a displacement vector that locates it in the next frame. Even though wider temporal context is freely available, prior efforts to take this into account have yielded only small…

2022

TIDEE: Tidying Up Novel Rooms Using Visuo-Semantic Commonsense Priors

ECCV 2022poster

"We introduce TIDEE, an embodied agent that tidies up a disordered scene based on learned commonsense object placement and room arrangement priors. TIDEE explores a home environment, detects objects that are out of their natural place, infers plausible object contexts for them, localizes such contex…

2021

CoCoNets: Continuous Contrastive 3D Scene Representations

CVPR 2021poster

This paper explores self-supervised learning of amodal 3D feature representations from RGB and RGB-D posed images and videos, agnostic to object and scene semantic content, and evaluates the resulting scene representations in the downstream tasks of visual correspondence, object tracking, and object…

Cited by 27PDFScholar
2021

Disentangling 3D Prototypical Networks for Few-Shot Concept Learning

ICLR 2021poster

We present neural architectures that disentangle RGB-D images into objects’ shapes and styles and a map of the background scene, and explore their applications for few-shot 3D object detection and few-shot concept classification. Our networks incorporate architectural biases that reflect the image f…

2021

Track, Check, Repeat: An EM Approach to Unsupervised Tracking

CVPR 2021poster

We propose an unsupervised method for detecting and tracking moving objects in 3D, in unlabelled RGB-D videos. The method begins with classic handcrafted techniques for segmenting objects using motion cues: we estimate optical flow and camera motion, and conservatively segment regions that appear to…

Cited by 9PDFScholar
2020

Embodied Language Grounding With 3D Visual Feature Representations

CVPR 2020poster

We propose associating language utterances to 3D visual abstractions of the scene they describe. The 3D visual abstractions are encoded as 3-dimensional visual feature maps. We infer these 3D visual scene feature maps from RGB images of the scene via view prediction: when the generated 3D scene feat…

Cited by 23PDFScholar
2020

Learning from Unlabelled Videos Using Contrastive Predictive Neural 3D Mapping

ICLR 2020poster

Predictive coding theories suggest that the brain learns by predicting observations at various levels of abstraction. One of the most basic prediction tasks is view prediction: how would a given scene look from an alternative viewpoint? Humans excel at this task. Our ability to imagine and fill in m…

Cited by 30SourcecodeScholar
2020

Tracking Emerges by Looking Around Static Scenes, with Neural 3D Mapping

ECCV 2020poster

with Neural 3D Mapping","We hypothesize that an agent that can look around in static scenes can learn rich visual representations applicable to 3D object tracking in complex dynamic scenes. We are motivated in this pursuit by the fact that the physical world itself is mostly static, and multiview co…

2018

Reward Learning From Narrated Demonstrations

CVPR 2018poster

Humans effortlessly “program” one another by communicating goals and desires in natural language. In contrast, humans program robotic behaviours by indicating desired object locations and poses to be achieved [5], by providing RGB images of goal configurations [19], or supplying a demonstration to b…

Cited by 46SourcePDFScholar
2017

Adversarial Inverse Graphics Networks: Learning 2D-To-3D Lifting and Image-To-Image Translation From Unpaired Supervision

ICCV 2017poster

Researchers have developed excellent feed-forward models that learn to map images to desired outputs, such as to the images' latent factors, or to other images, using supervised learning. Learning such mappings from unlabelled data, or improving upon supervised models by exploiting unlabelled data,…

Cited by 170PDFScholar
2017

Segmentation-Aware Convolutional Networks Using Local Attention Masks

ICCV 2017poster

We introduce an approach to integrate segmentation information within a convolutional neural network (CNN). This counter-acts the tendency of CNNs to smooth information across regions and increases their spatial precision. To obtain segmentation information, we set up a CNN to provide an embedding s…

Cited by 188PDFcodeScholar